QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer Inference
Xinkuang Geng, Siting Liu, Leibo Liu, Jie Han, Honglan Jiang
摘要
While exhibiting superior performance in many tasks, vision transformers (ViTs) face challenges in quantization. Some existing low-bit-width quantization techniques cannot effectively cover the whole inference process of ViTs, leading to an additional memory overhead (22.3%-172.6%) compared with corresponding fully quantized models. To address this issue, we propose quadruplet uniform quantization (QUQ) to deal with data of various distributions in ViT. QUQ divides the entire data range into at most four subranges that are uniformly quantized with different scale factors. To determine the partition scheme and quantization parameters, an efficient relaxation algorithm is proposed accordingly. Moreover, dedicated encoding and decoding strategies are devised to facilitate the design of an efficient accelerator. Experimental results show that QUQ surpasses state-of-the-art quantization techniques; it is the first viable scheme that can fully quantize ViTs to 6-bit with acceptable accuracy. Compared with conventional uniform quantization, QUQ leads to not only a higher accuracy but also an accelerator with lower area and power.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Accurate Post-Training Quantization for Vision TransformerYifu Ding, Haotong Qin, Qinghua Yan, Zhenhua Chai 等ACM MM 2022 · 被引用 68 次
相关 Paper
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 被引用 10 次
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 被引用 172 次
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao 等NeurIPS 2022 · 被引用 185 次
- LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision TransformersMinjun Kim, Jaeri Lee, Jongjin Kim, Jeongin Yun 等AAAI 2026 · 被引用 1 次
- PackQViT: Faster Sub-8-bit Vision Transformers via Full and Packed Quantization on the MobilePeiyan Dong, Lei Lu, Chao Wu, Cheng Lyu 等NeurIPS 2023 · 被引用 44 次
