Bit-shrinking: Limiting Instantaneous Sharpness for Improving Post-training Quantization
Chen Lin, Bo Peng, Zheyang Li, Wenming Tan, Ye Ren, Jun Xiao, Shiliang Pu
Abstract
Post-training quantization (PTQ) is an effective compression method to reduce the model size and computational cost. However, quantizing a model into a low-bit one, e.g., lower than 4, is difficult and often results in nonnegligible performance degradation. To address this, we investigate the loss landscapes of quantized networks with various bit-widths. We show that the network with more ragged loss surface, is more easily trapped into bad local minima, which mostly appears in low-bit quantization. A deeper analysis indicates, the ragged surface is caused by the injection of excessive quantization noise. To this end, we detach a sharpness term from the loss which reflects the impact of quantization noise. To smooth the rugged loss surface, we propose to limit the sharpness term small and stable during optimization. Instead of directly optimizing the target bit network, we design a self-adapted shrinking scheduler for the bit-width in continuous domain from high bit-width to the target by limiting the increasing sharpness term within a proper range. It can be viewed as iteratively adding small "instant" quantization noise and adjusting the network to eliminate its impact. Widely experiments including classification and detection tasks demonstrate the effectiveness of the Bit-shrinking strategy in PTQ. On the Vision Transformer models, our INT8 and INT6 models drop within 0.5% and 1.5% Top-1 accuracy, respectively. On the traditional CNN networks, our INT4 quantized models drop within 1.3% and 3.5% Top-1 accuracy on ResNet18 and MobileNetV2 without fine-tuning, which achieves the state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6cf55ef8-9a8f-4ae5-9b45-d9734d7e5decCited by top-tier papers10
- MagR: Weight Magnitude Reduction for Enhancing Post-Training QuantizationAozhong Zhang, Naigang Wang, Yanxia Deng, Xin Li et al.NeurIPS 2024 · 33 citations
- Outlier-aware Slicing for Post-Training Quantization in Vision TransformerYuexiao Ma, Huixia Li, Xiawu Zheng, Feng Ling et al.ICML 2024 · 17 citations
- ERQ: Error Reduction for Post-Training Quantization of Vision TransformersYunshan Zhong, Jiawei Hu, You Huang, Yuxin Zhang et al.ICML 2024 · 14 citations
- Sharpness-Aware Data Generation for Zero-shot QuantizationHoang Anh Dung, Cuong Pham, Trung Le, Jianfei Cai et al.ICML 2024 · 8 citations
- Gradient-Aligned Calibration for Post-Training Quantization of Diffusion ModelsDung Anh Hoang, Cuong Pham, Trung Le, Jianfei Cai et al.ICLR 2026 · 1 citation
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- PD-Quant: Post-Training Quantization Based on Prediction Difference MetricJiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang et al.CVPR 2023
- Towards Accurate Post-training Network Quantization via Bit-Split and StitchingPeisong Wang, Qiang Chen, Xiangyu He, Jian ChengICML 2020 · 159 citations
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang et al.NeurIPS 2021 · 528 citations
- Mr.BiQ: Post-Training Non-Uniform Quantization based on Minimizing the Reconstruction ErrorYongkweon Jeon, Chungman Lee, Eulrang Cho, Yeonju RoCVPR 2022 · 28 citations
- QSCA: Quantization with Self-Compensating Auxiliary for Monocular Depth EstimationJincheol Yang, Jaemin Choi, Matti Zinke, Suk-Ju KangNeurIPS 2025 · 1 citation
