Rethinking Differentiable Search for Mixed-Precision Neural Networks
Zhaowei Cai, Nuno Vasconcelos
Abstract
Low-precision networks, with weights and activations quantized to low bit-width, are widely used to accelerate inference on edge devices. However, current solutions are uniform, using identical bit-width for all filters. This fails to account for the different sensitivities of different filters and is suboptimal. Mixed-precision networks address this problem, by tuning the bit-width to individual filter requirements. In this work, the problem of optimal mixed-precision network search (MPS) is considered. To circumvent its difficulties of discrete search space and combinatorial optimization, a new differentiable search architecture is proposed, with several novel contributions to advance the efficiency by leveraging the unique properties of the MPS problem. The resulting Efficient differentiable MIxed-Precision network Search (EdMIPS) method is effective at finding the optimal bit allocation for multiple popular networks, and can search a large model, e.g. Inception-V3, directly on Im-ageNet without proxy task in a reasonable amount of time. The learned mixed-precision networks significantly outperform their uniform counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01f23651-ab4a-49b1-b09f-60493f2f98d6Cited by top-tier papers33
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationCong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng et al.ISCA 2023 · 151 citations
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu et al.MICRO 2022 · 109 citations
- DominoSearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networksWei Sun, Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven et al.NeurIPS 2021 · 67 citations
- Wavelet Feature Maps Compression for Image-to-Image CNNsShahaf E. Finder, Yair Zohav, Maor Ashkenazi, Eran TreisterNeurIPS 2022 · 63 citations
- Dynamic Network Quantization for Efficient Video InferenceXimeng Sun, Rameswar Panda, Chun-Fu (Richard) Chen, Aude Oliva et al.ICCV 2021 · 56 citations
Builds on1
Related papers
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu et al.ICML 2022 · 49 citations
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama et al.ICLR 2020 · 159 citations
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 95 citations
- No Retraining at Edge: Efficient Resource-Aware Mixed-Precision Quantization via Federated Supernet LearningLianbo Ma, Yonghui Su, Nan Li, Xingwei WangICML 2026
