PD-Quant: Post-Training Quantization Based on Prediction Difference Metric
Jiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang, Xinggang Wang, Wenyu Liu
Abstract
Post-training quantization (PTQ) is a neural network compression technique that converts a full-precision model into a quantized model using lower-precision data types. Although it can help reduce the size and computational cost of deep neural networks, it can also introduce quantization noise and reduce prediction accuracy, especially in extremely low-bit settings. How to determine the appropriate quantization parameters (e.g., scaling factors and rounding of weights) is the main problem facing now. Existing methods attempt to determine these parameters by minimize the distance between features before and after quantization, but such an approach only considers local information and may not result in the most optimal quantization parameters. We analyze this issue and propose PD-Quant, a method that addresses this limitation by considering global information. It determines the quantization parameters by using the information of differences between network prediction before and after quantization. In addition, PD-Quant can alleviate the overfitting problem in PTQ caused by the small number of calibration sets by adjusting the distribution of activations. Experiments show that PD-Quant leads to better quantization parameters and improves the prediction accuracy of quantized models, especially in low-bit settings. For example, PD-Quant pushes the accuracy of ResNet-18 up to 53.14% and RegNetX-600MF up to 40.67% in weight 2-bit activation 2-bit. The code is released at https://github.com/hustvl/PD-Quant .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53a6d73d-d870-4ec9-b591-ca675062c28cCited by top-tier papers39
- PTQ4DiT: Post-training Quantization for Diffusion TransformersJunyi Wu, Haoxuan Wang, Yuzhang Shang, Mubarak Shah et al.NeurIPS 2024 · 87 citations
- TinySAM: Pushing the Envelope for Efficient Segment Anything ModelHan Shu, Wenshuo Li, Yehui Tang, Yiman Zhang et al.AAAI 2025 · 57 citations
- LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object DetectionSifan Zhou, Liang Li, Xinyu Zhang, Bo Zhang et al.ICLR 2024 · 40 citations
- PTQ4SAM: Post-Training Quantization for Segment AnythingChengtao Lv, Hong Chen, Jinyang Guo, Yifu Ding et al.CVPR 2024 · 22 citations
- Purifying Quantization-conditioned Backdoors via Layer-wise Activation Correction with Distribution ApproximationBoheng Li, Yishuo Cai, Jisong Cai, Yiming Li et al.ICML 2024 · 19 citations
Builds on14
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training QuantizationXiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu et al.ICLR 2022 · 248 citations
- Accurate Post Training Quantization With Small Calibration SetsItay Hubara, Yury Nahshan, Yair Hanani, Ron Banner et al.ICML 2021 · 238 citations
Related papers
- Leveraging Inter-Layer Dependency for Post -Training QuantizationChangbao Wang, Dandan Zheng, Yuanliu Liu, Liang LiNeurIPS 2022 · 27 citations
- Bit-shrinking: Limiting Instantaneous Sharpness for Improving Post-training QuantizationChen Lin, Bo Peng, Zheyang Li, Wenming Tan et al.CVPR 2023
- RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersZhikai Li, Junrui Xiao, Lianwei Yang, Qingyi GuICCV 2023 · 172 citations
- Enhancing Post-Training Quantization Calibration Through Contrastive LearningYuzhang Shang, Gaowen Liu, Ramana Rao Kompella, Yan YanCVPR 2024 · 10 citations
- Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank CompensationZhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn et al.AAAI 2024 · 50 citations
