Retraining-free Model Quantization via One-Shot Weight-Coupling Learning
Chen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie, Rongwei Lu, Xinzhu Ma, Zhi Wang, Wenwu Zhu
Abstract
Quantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suf-fers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quan-tization (MPQ) is advocated to compress the model ef-fectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage effi-ciently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency sig-nificantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultane-ously within a set of shared weights. However, our ob-servations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dy-namically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged prop-erly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the good-ness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effective-ness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47c5a7ff-0031-444b-84be-d497f96fbb40Cited by top-tier papers7
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model AccelerationYe Li, Yuan Meng, Zewen Sun, Kangye Ji et al.ICLR 2026 · 60 citations
- Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset SamplingJinhee Kim, Jae Jun An, Kang Eun Jeon, Jong Hwan KoNeurIPS 2025 · 4 citations
- Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness PerspectiveJiacheng Jiang, Yuan Meng, Chen Tang, Han Yu et al.ACM MM 2025 · 1 citation
- No Retraining at Edge: Efficient Resource-Aware Mixed-Precision Quantization via Federated Supernet LearningLianbo Ma, Yonghui Su, Nan Li, Xingwei WangICML 2026
- SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer ProgrammingShuzhao Xie, Jiahang Liu, Weixiang Zhang, Shijia Ge et al.ACM MM 2025
Builds on30
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
Related papers
- InfoQ: Mixed-Precision Quantization via Global Information FlowMehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Manuel RoveriAAAI 2026 · 2 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
- One-Shot Model for Mixed-Precision QuantizationIvan Koryakovskiy, Alexandra Yakovleva, Valentin Buchnev, Temur Isaev et al.CVPR 2023
- Double Rounding: Nearly Lossless Adaptive Bit Switching for QATHaiduo Huang, Zhenhua Liu, Tian Xia, Pengju RenAAAI 2026
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
