RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision Transformers
Zhikai Li, Junrui Xiao, Lianwei Yang, Qingyi Gu
Abstract
Post-training quantization (PTQ), which only requires a tiny dataset for calibration without end-to-end retraining, is a light and practical model compression technique. Recently, several PTQ schemes for vision transformers (ViTs) have been presented; unfortunately, they typically suffer from non-trivial accuracy degradation, especially in low-bit cases. In this paper, we propose RepQ-ViT, a novel PTQ framework for ViTs based on quantization scale reparameterization, to address the above issues. RepQ-ViT decouples the quantization and inference processes, where the former employs complex quantizers and the latter employs scale-reparameterized simplified quantizers. This ensures both accurate quantization and efficient inference, which distinguishes it from existing approaches that sacrifice quantization performance to meet the target hardware. More specifically, we focus on two components with extreme distributions: post-LayerNorm activations with severe inter-channel variation and post-Softmax activations with power-law features, and initially apply channel-wise quantization and log quantization, respectively. Then, we reparameterize the scales to hardware-friendly layer-wise quantization and log2 quantization for inference, with only slight accuracy or computational costs. Extensive experiments are conducted on multiple vision tasks with different model variants, proving that RepQ-ViT, without hyperparameters and expensive reconstruction procedures, can outperform existing strong baselines and encouragingly improve the accuracy of 4-bit PTQ of ViTs to a usable level. Code is available at https://github.com/zkkli/RepQ-ViT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3663321-320d-41d3-8f4c-e4f0b4df856fCited by top-tier papers65
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
- I-ViT: Integer-only Quantization for Efficient Vision Transformer InferenceZhikai Li, Qingyi GuICCV 2023 · 176 citations
- PTQ4DiT: Post-training Quantization for Diffusion TransformersJunyi Wu, Haoxuan Wang, Yuzhang Shang, Mubarak Shah et al.NeurIPS 2024 · 87 citations
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 40 citations
- PTQ4SAM: Post-Training Quantization for Segment AnythingChengtao Lv, Hong Chen, Jinyang Guo, Yifu Ding et al.CVPR 2024 · 22 citations
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
Related papers
- Towards Accurate Post-Training Quantization for Vision TransformerYifu Ding, Haotong Qin, Qinghua Yan, Zhenhua Chai et al.ACM MM 2022 · 68 citations
- AIQViT: Architecture-Informed Post-Training Quantization for Vision TransformersRunqing Jiang, Ye Zhang, Longguang Wang, Pengpeng Yu et al.AAAI 2025 · 4 citations
- Instance-Aware Group Quantization for Vision TransformersJaehyeon Moon, Dohyung Kim, Junyong Cheon, Bumsub HamCVPR 2024 · 11 citations
- GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersGuang Liang, Xinyao Liu, Jianxin WuNeurIPS 2025 · 10 citations
- FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix ApproximationZhuguanyu Wu, Shihe Wang, Jiayi Zhang, Jiaxin Chen et al.CVPR 2025
