Generalizable Mixed-Precision Quantization via Attribution Rank Preservation
Ziwei Wang, Han Xiao, Jiwen Lu, Jie Zhou
Abstract
In this paper, we propose a generalizable mixed-precision quantization (GMPQ) method for efficient inference. Conventional methods require the consistency of datasets for bitwidth search and model deployment to guarantee the policy optimality, leading to heavy search cost on challenging largescale datasets in realistic applications. On the contrary, our GMPQ searches the mixed-quantization policy that can be generalized to largescale datasets with only a small amount of data, so that the search cost is significantly reduced without performance degradation. Specifically, we observe that locating network attribution correctly is general ability for accurate visual analysis across different data distribution. Therefore, despite of pursuing higher model accuracy and complexity, we preserve attribution rank consistency between the quantized models and their full-precision counterparts via efficient capacity-aware attribution imitation for generalizable mixed-precision quantization strategy search. Extensive experiments show that our method obtains competitive accuracy-complexity trade-off compared with the state-of-the-art mixed-precision networks in significantly reduced search cost. The code is available at https://github.com/ZiweiWangTHU/GMPQ.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin RegularizationChen Tang, Kai Ouyang, Zenghao Chai, Yunpeng Bai et al.ACM MM 2023 · 11 citations
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie et al.CVPR 2024 · 5 citations
- Efficient and Generalizable Mixed-Precision Quantization via Topological EntropyNan Li, Yonghui Su, Lianbo MaNeurIPS 2025 · 4 citations
- One-Shot Model for Mixed-Precision QuantizationIvan Koryakovskiy, Alexandra Yakovleva, Valentin Buchnev, Temur Isaev et al.CVPR 2023
- Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient AligningLianbo Ma, Jianlun Ma, Yuee Zhou, Guoyang Xie et al.ICML 2025
Builds on12
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami et al.NeurIPS 2020 · 434 citations
- ThunderNet: Towards Real-Time Generic Object Detection on Mobile DevicesZheng Qin, Zeming Li, Zhaoning Zhang, Yiping Bao et al.ICCV 2019 · 282 citations
- Towards Accurate Post-training Network Quantization via Bit-Split and StitchingPeisong Wang, Qiang Chen, Xiangyu He, Jian ChengICML 2020 · 159 citations
Related papers
- InfoQ: Mixed-Precision Quantization via Global Information FlowMehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Manuel RoveriAAAI 2026 · 2 citations
- Rethinking Differentiable Search for Mixed-Precision Neural NetworksZhaowei Cai, Nuno VasconcelosCVPR 2020
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language ModelsPengxiang Zhao, Xiaoming YuanICML 2025
- Once Quantization-Aware Training: High Performance Extremely Low-bit Architecture SearchMingzhu Shen, Feng Liang, Ruihao Gong, Yuhang Li et al.ICCV 2021 · 50 citations
- Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic ProgrammingZihao Deng, Sayeh Sharify, Xin Wang, Michael OrshanskyDAC 2025 · 1 citation
