Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
Xingchao Liu, Mao Ye, Dengyong Zhou, Qiang Liu
Abstract
We consider the post-training quantization problem, which discretizes the weights of pre-trained deep neural networks without re-training the model. We propose multipoint quantization, a quantization method that approximates a fullprecision weight vector using a linear combination of multiple vectors of low-bit numbers; this is in contrast to typical quantization methods that approximate each weight using a single low precision number. Computationally, we construct the multipoint quantization with an efficient greedy selection procedure, and adaptively decides the number of low precision points on each quantized weight vector based on the error of its output. This allows us to achieve higher precision levels for important weights that greatly influence the outputs, yielding an "effect of mixed precision" but without physical mixed precision implementations (which requires specialized hardware accelerators (Wang et al. 2019) ). Empirically, our method can be implemented by common operands, bringing almost no memory and computation overhead. We show that our method outperforms a range of state-of-the-art methods on ImageNet classification and it can be generalized to more challenging tasks like PASCAL VOC object detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationYunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu et al.CVPR 2022 · 79 citations
- A Langevin-like Sampler for Discrete DistributionsRuqi Zhang, Xingchao Liu, Qiang LiuICML 2022 · 51 citations
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu et al.NeurIPS 2023 · 24 citations
- Be Like Water: Adaptive Floating Point for Machine LearningThomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang et al.ICML 2022 · 11 citations
- Efficient Transformer-based 3D Object Detection with Dynamic Token HaltingMao Ye, Gregory P. Meyer, Yuning Chai, Qiang LiuICCV 2023 · 10 citations
Builds on2
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 622 citations
Related papers
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- Leveraging Inter-Layer Dependency for Post -Training QuantizationChangbao Wang, Dandan Zheng, Yuanliu Liu, Liang LiNeurIPS 2022 · 27 citations
- PD-Quant: Post-Training Quantization Based on Prediction Difference MetricJiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang et al.CVPR 2023
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 17 citations
- Rethinking Asymmetric Quantization: Hidden Symmetry in Vision Model WeightsMasafumi Mori, Shinya Gongyo, Mitsuru AmbaiCVPR 2026
