AlignQ: Alignment Quantization with ADMM-based Correlation Preservation
Ting-An Chen, De-Nian Yang, Ming-Syan Chen
Abstract
Quantization is an efficient network compression approach to reduce the inference time. However, existing approaches ignored the distribution difference between training and testing data, thereby inducing a large quantization error in inference. To address this issue, we propose a new quantization scheme, Alignment Quantization with ADMMbased Correlation Preservation (AlignQ), which exploits the cumulative distribution function (CDF) to align the data to be i.i.d. (independently and identically distributed) for quantization error minimization. Afterward, our theoretical analysis indicates that the significant changes in data correlations after the quantization induce a large quantization error. Accordingly, we aim to preserve the relationship of data from the original space to the aligned quantization space for retaining the prediction information. We design an optimization process by leveraging the Alternating Direction Method of Multipliers (ADMM) optimization to minimize the differences in data correlations before and after the alignment and quantization. In experiments, we visualize non-i.i.d. in training and testing data in the benchmark. We further adopt domain shift data to compare AlignQ with the state-of-the-art. Experimental results show that AlignQ achieves significant performance improvements especially in low-bit models. Code is available at https: //github.com/tinganchen/AlignQ.git . Quantization Large quantization error Training data Testing data Training space Testing space Aligned space for Quantization -0.4 -0.6 +0.8 (c) Data correlations preservation (a) Non-i.i.d data (b) Alignment
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae08e809-d93d-4bed-a7a6-ecc512aaeb5bCited by top-tier papers3
- ClimbQ: Class Imbalanced Quantization Enabling Robustness on Efficient InferencesTing-An Chen, De-Nian Yang, Ming-Syan ChenNeurIPS 2022 · 8 citations
- Overcoming Forgetting Catastrophe in Quantization-Aware TrainingTing-An Chen, De-Nian Yang, Ming-Syan ChenICCV 2023 · 4 citations
- Data-Free Quantization via Pseudo-label FilteringChunxiao Fan, Ziqi Wang, Dan Guo, Meng WangCVPR 2024
Builds on5
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 315 citations
- Linear Symmetric Quantization of Neural Networks for Low-precision Integer HardwareXiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu et al.ICLR 2020 · 67 citations
- ZeroQ: A Novel Zero Shot Quantization FrameworkYaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami et al.CVPR 2020
- Zero-Shot Adversarial QuantizationYuang Liu, Wei Zhang, Jun WangCVPR 2021
Related papers
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham et al.ICLR 2020 · 157 citations
- Fixed-Point Back-Propagation TrainingXishan Zhang, Shaoli Liu, Rui Zhang, Chang Liu et al.CVPR 2020
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- Distribution-Aware Adaptive Multi-Bit QuantizationSijie Zhao, Tao Yue, Xuemei HuCVPR 2021
- Diversifying Sample Generation for Accurate Data-Free QuantizationXiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong et al.CVPR 2021
