Finding the Task-Optimal Low-Bit Sub-Distribution in Deep Neural Networks
Runpei Dong, Zhanhong Tan, Mengdi Wu, Linfeng Zhang, Kaisheng Ma
Abstract
Quantized neural networks typically require smaller memory footprints and lower computation complexity, which is crucial for efficient deployment. However, quantization inevitably leads to a distribution divergence from the original network, which generally degrades the performance. To tackle this issue, massive efforts have been made, but most existing approaches lack statistical considerations and depend on several manual configurations. In this paper, we present an adaptive-mapping quantization method to learn an optimal latent sub-distribution that is inherent within models and smoothly approximated with a concrete Gaussian Mixture (GM). In particular, the network weights are projected in compliance with the GM-approximated sub-distribution. This sub-distribution evolves along with the weight update in a co-tuning schema guided by the direct task-objective optimization. Sufficient experiments on image classification and object detection over various modern architectures demonstrate the effectiveness, generalization property, and transferability of the proposed method. Besides, an efficient deployment flow for the mobile CPU is developed, achieving up to 7.46 inference acceleration on an octa-core ARM CPU. Our codes have been publicly released at https://github.com/RunpeiDong/DGMS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi et al.ICLR 2024 · 315 citations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- Compressing LLMs: The Truth is Rarely Pure and Never SimpleAjay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang et al.ICLR 2024 · 61 citations
- VPP: Efficient Conditional 3D Generation via Voxel-Point Progressive RepresentationZekun Qi, Muzhou Yu, Runpei Dong, Kaisheng MaNeurIPS 2023 · 22 citations
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 10 citations
Builds on10
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami et al.ICML 2021 · 240 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad et al.NeurIPS 2020 · 149 citations
Related papers
- Network Memory Footprint Compression Through Jointly Learnable Codebooks and MappingsEdouard Yvinec, Arnaud Dapogny, Kevin BaillyICLR 2024 · 2 citations
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language ModelsPengxiang Zhao, Xiaoming YuanICML 2025
- AlignQ: Alignment Quantization with ADMM-based Correlation PreservationTing-An Chen, De-Nian Yang, Ming-Syan ChenCVPR 2022 · 4 citations
- Towards Accurate Low Bit-Width Quantization with Multiple Phase AdaptationsZhaoyi Yan, Yemin Shi, Yaowei Wang, Mingkui Tan et al.AAAI 2020 · 2 citations
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
