Mixed Precision DNNs: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso García, Stephen Tiedemann, Thomas Kemp, Akira Nakamura
Abstract
Efficient deep neural network (DNN) inference on mobile or embedded devices typically involves quantization of the network parameters and activations. In particular, mixed precision networks achieve better performance than networks with homogeneous bitwidth for the same size constraint. Since choosing the optimal bitwidths is not straight forward, training methods, which can learn them, are desirable. Differentiable quantization with straight-through gradients allows to learn the quantizer's parameters using gradient methods. We show that a suited parametrization of the quantizer is the key to achieve a stable training and a good final performance. Specifically, we propose to parametrize the quantizer with the step size and dynamic range. The bitwidth can then be inferred from them. Other parametrizations, which explicitly use the bitwidth, consistently perform worse. We confirm our findings with experiments on CIFAR-10 and ImageNet and we obtain mixed precision DNNs with learned quantization parameters, achieving state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fedc9d56-c1bf-4b61-847a-1f3f2362fd11Cited by top-tier papers33
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Overcoming Oscillations in Quantization-Aware TrainingMarkus Nagel, Marios Fournarakis, Yelysei Bondarenko, Tijmen BlankevoortICML 2022 · 163 citations
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel et al.NeurIPS 2022 · 154 citations
- Pruning vs Quantization: Which is Better?Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi et al.NeurIPS 2023 · 152 citations
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad et al.NeurIPS 2020 · 149 citations
Builds on1
Related papers
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 95 citations
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive ResolutionZhaoyang Zhang, Wenqi Shao, Jinwei Gu, Xiaogang Wang et al.ICML 2021 · 36 citations
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu et al.ICML 2022 · 49 citations
- Not All Bits have Equal Value: Heterogeneous Precisions via Trainable NoisePedro Savarese, Xin Yuan, Yanjing Li, Michael MaireNeurIPS 2022 · 9 citations
