QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Xiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu, Fengwei Yu
Abstract
Recently, post-training quantization (PTQ) has driven much attention to produce efficient neural networks without long-time retraining. Despite its low cost, current PTQ works tend to fail under the extremely low-bit setting. In this study, we pioneeringly confirm that properly incorporating activation quantization into the PTQ reconstruction benefits the final accuracy. To deeply understand the inherent reason, a theoretical framework is established, indicating that the flatness of the optimized low-bit model on calibration and test data is crucial. Based on the conclusion, a simple yet effective approach dubbed as QDROP is proposed, which randomly drops the quantization of activations during PTQ. Extensive experiments on various tasks including computer vision (image classification, object detection) and natural language processing (text classification and question answering) prove its superiority. With QDROP, the limit of PTQ is pushed to the 2-bit activation for the first time and the accuracy boost can be up to 51.49%. Without bells and whistles, QDROP establishes a new state of the art for PTQ. Our code is available at https://github.com/wimh966/QDrop and has been integrated into MQBench ( https://github.com/ModelTC/MQBench ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers81
- Outlier Suppression: Pushing the Limit of Low-bit Transformer Language ModelsXiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong et al.NeurIPS 2022 · 238 citations
- PTQD: Accurate Post-Training Quantization for Diffusion ModelsYefei He, Luping Liu, Jing Liu, Weijia Wu et al.NeurIPS 2023 · 219 citations
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 172 citations
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim et al.NeurIPS 2023 · 109 citations
- PTQ4DiT: Post-training Quantization for Diffusion TransformersJunyi Wu, Haoxuan Wang, Yuzhang Shang, Mubarak Shah et al.NeurIPS 2024 · 87 citations
Builds on19
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
Related papers
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- PD-Quant: Post-Training Quantization Based on Prediction Difference MetricJiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang et al.CVPR 2023
- Leveraging Inter-Layer Dependency for Post -Training QuantizationChangbao Wang, Dandan Zheng, Yuanliu Liu, Liang LiNeurIPS 2022 · 27 citations
- QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language ModelsJing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong et al.ICLR 2024 · 75 citations
- PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang et al.ACL 2025 · 7 citations
