AlignQ: Alignment Quantization with ADMM-based Correlation Preservation
Ting-An Chen, De-Nian Yang, Ming-Syan Chen
摘要
Quantization is an efficient network compression approach to reduce the inference time. However, existing approaches ignored the distribution difference between training and testing data, thereby inducing a large quantization error in inference. To address this issue, we propose a new quantization scheme, Alignment Quantization with ADMMbased Correlation Preservation (AlignQ), which exploits the cumulative distribution function (CDF) to align the data to be i.i.d. (independently and identically distributed) for quantization error minimization. Afterward, our theoretical analysis indicates that the significant changes in data correlations after the quantization induce a large quantization error. Accordingly, we aim to preserve the relationship of data from the original space to the aligned quantization space for retaining the prediction information. We design an optimization process by leveraging the Alternating Direction Method of Multipliers (ADMM) optimization to minimize the differences in data correlations before and after the alignment and quantization. In experiments, we visualize non-i.i.d. in training and testing data in the benchmark. We further adopt domain shift data to compare AlignQ with the state-of-the-art. Experimental results show that AlignQ achieves significant performance improvements especially in low-bit models. Code is available at https: //github.com/tinganchen/AlignQ.git . Quantization Large quantization error Training data Testing data Training space Testing space Aligned space for Quantization -0.4 -0.6 +0.8 (c) Data correlations preservation (a) Non-i.i.d data (b) Alignment
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ClimbQ: Class Imbalanced Quantization Enabling Robustness on Efficient InferencesTing-An Chen, De-Nian Yang, Ming-Syan ChenNeurIPS 2022 · 被引用 8 次
- Overcoming Forgetting Catastrophe in Quantization-Aware TrainingTing-An Chen, De-Nian Yang, Ming-Syan ChenICCV 2023 · 被引用 4 次
- Data-Free Quantization via Pseudo-label FilteringChunxiao Fan, Ziqi Wang, Dan Guo, Meng WangCVPR 2024
它引用的顶会 Paper5
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 被引用 315 次
- Linear Symmetric Quantization of Neural Networks for Low-precision Integer HardwareXiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu 等ICLR 2020 · 被引用 67 次
- ZeroQ: A Novel Zero Shot Quantization FrameworkYaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami 等CVPR 2020
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
相关 Paper
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham 等ICLR 2020 · 被引用 157 次
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu 等CVPR 2020
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- Distribution-Aware Adaptive Multi-Bit QuantizationSijie Zhao, Tao Yue, Xuemei HuCVPR 2021
- Diversifying Sample Generation for Accurate Data-Free QuantizationXiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong 等CVPR 2021
