On the Impact of Weight Quantization on Deep Neural Network Uncertainty
Shuang Liang, Xun Lu, Zi-Ang Liu, Ming-Liang Wang, Yan Lyu, Shao-Qun Zhang
Abstract
Weight Quantization (WQ) is a key technique for lightweight Deep Neural Network (DNN) computations. While existing algorithms often pursue memory compression and inference acceleration with accuracy comparable to full-precision models, the effect of WQ on DNN uncertainty remains largely unexplored. In this paper, we quantify the impact of WQ on DNN uncertainty through the novel Exact Moment Propagation (EMP) uncertainty estimator. It is observed that WQ significantly increases DNN uncertainty. Based on the EMP estimator, we propose the MOMent Alignment (MOMA) to reduce WQ-induced uncertainty and preserve the accuracy of weight-quantized DNNs. Empirical results across various DNN architectures and datasets validate the effectiveness of both EMP and MOMA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5010313-ba11-4b3e-bdc8-6717aa088bfaBuilds on5
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao et al.NeurIPS 2022 · 185 citations
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- Training Uncertainty-Aware Classifiers with Conformalized Deep LearningBat-Sheva Einbinder, Yaniv Romano, Matteo Sesia, Yanfei ZhouNeurIPS 2022 · 84 citations
- Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the EdgeXuan Shen, Peiyan Dong, Lei Lu, Zhenglun Kong et al.AAAI 2024 · 59 citations
- Training Binary Neural Networks using the Bayesian Learning RuleXiangming Meng, Roman Bachmann, Mohammad Emtiyaz KhanICML 2020 · 47 citations
Related papers
- Lightweight Approaches to DNN Regression Error Reduction: An Uncertainty Alignment PerspectiveZenan Li, Maorun Zhang, Jingwei Xu, Yuan Yao et al.ICSE 2023 · 3 citations
- EMPIR: Ensembles of Mixed Precision Deep Networks for Increased Robustness Against Adversarial AttacksSanchari Sen, Balaraman Ravindran, Anand RaghunathanICLR 2020 · 69 citations
- Towards Certificated Model Robustness Against Weight PerturbationsTsui-Wei Weng, Pu Zhao, Sijia Liu, Pin-Yu Chen et al.AAAI 2020 · 33 citations
- INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair EncodingFangxin Liu, Ning Yang, Zhiyan Song, Zongwu Wang et al.DAC 2024 · 10 citations
- SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian ApproximationCong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao et al.ICLR 2022 · 92 citations
