Q-ViT: Accurate and Fully Quantized Low-bit Vision Transformer
Yanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao, Peng Gao, Guodong Guo
摘要
The large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14× and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language ModelsJing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong 等ICLR 2024 · 被引用 75 次
- Q-DM: An Efficient Low-bit Quantized Diffusion ModelYanjing Li, Sheng Xu, Xianbin Cao, Xiao Sun 等NeurIPS 2023 · 被引用 70 次
- Oscillation-free Quantization for Low-bit Vision TransformersShih-Yang Liu, Zechun Liu, Kwang-Ting ChengICML 2023 · 被引用 63 次
- TinySAM: Pushing the Envelope for Efficient Segment Anything ModelHan Shu, Wenshuo Li, Yehui Tang, Yiman Zhang 等AAAI 2025 · 被引用 57 次
它引用的顶会 Paper15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
相关 Paper
- Bi-ViT: Pushing the Limit of Vision Transformer QuantizationYanjing Li, Sheng Xu, Mingbao Lin, Xianbin Cao 等AAAI 2024 · 被引用 23 次
- MiniViT: Compressing Vision Transformers with Weight MultiplexingJinnian Zhang, Houwen Peng, Kan Wu, Mengchen Liu 等CVPR 2022 · 被引用 115 次
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 被引用 21 次
- QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceXinkuang Geng, Siting Liu, Leibo Liu, Jie Han 等DAC 2024 · 被引用 5 次
- Instance-Aware Group Quantization for Vision TransformersJaehyeon Moon, Dohyung Kim, Junyong Cheon, Bumsub HamCVPR 2024 · 被引用 11 次
