Post-training Quantization with Multiple Points: Mixed Precision without Mixed Precision
Xingchao Liu, Mao Ye, Dengyong Zhou, Qiang Liu
摘要
We consider the post-training quantization problem, which discretizes the weights of pre-trained deep neural networks without re-training the model. We propose multipoint quantization, a quantization method that approximates a fullprecision weight vector using a linear combination of multiple vectors of low-bit numbers; this is in contrast to typical quantization methods that approximate each weight using a single low precision number. Computationally, we construct the multipoint quantization with an efficient greedy selection procedure, and adaptively decides the number of low precision points on each quantized weight vector based on the error of its output. This allows us to achieve higher precision levels for important weights that greatly influence the outputs, yielding an "effect of mixed precision" but without physical mixed precision implementations (which requires specialized hardware accelerators (Wang et al. 2019) ). Empirically, our method can be implemented by common operands, bringing almost no memory and computation overhead. We show that our method outperforms a range of state-of-the-art methods on ImageNet classification and it can be generalized to more challenging tasks like PASCAL VOC object detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationYunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu 等CVPR 2022 · 被引用 79 次
- A Langevin-like Sampler for Discrete DistributionsRuqi Zhang, Xingchao Liu, Qiang LiuICML 2022 · 被引用 51 次
- TexQ: Zero-shot Network Quantization with Texture Feature Distribution CalibrationXinrui Chen, Yizhi Wang, Renao Yan, Yiqing Liu 等NeurIPS 2023 · 被引用 24 次
- Be Like Water: Adaptive Floating Point for Machine LearningThomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang 等ICML 2022 · 被引用 11 次
- Efficient Transformer-based 3D Object Detection with Dynamic Token HaltingMao Ye, Gregory P. Meyer, Yuning Chai, Qiang LiuICCV 2023 · 被引用 10 次
它引用的顶会 Paper2
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
相关 Paper
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang 等ICLR 2021 · 被引用 619 次
- Leveraging Inter-Layer Dependency for Post -Training QuantizationChangbao Wang, Dandan Zheng, Yuanliu Liu, Liang LiNeurIPS 2022 · 被引用 27 次
- PD-Quant: Post-Training Quantization Based on Prediction Difference MetricJiawei Liu, Lin Niu, Zhihang Yuan, Dawei Yang 等CVPR 2023
- Multi-Precision Policy Enforced Training (MuPPET) : A Precision-Switching Strategy for Quantised Fixed-Point Training of CNNsAditya Rajagopal, Diederik Adriaan Vink, Stylianos I. Venieris, Christos-Savvas BouganisICML 2020 · 被引用 17 次
- Rethinking Asymmetric Quantization: Hidden Symmetry in Vision Model WeightsMasafumi Mori, Shinya Gongyo, Mitsuru AmbaiCVPR 2026
