Mixed Precision DNNs: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso García, Stephen Tiedemann, Thomas Kemp, Akira Nakamura
摘要
Efficient deep neural network (DNN) inference on mobile or embedded devices typically involves quantization of the network parameters and activations. In particular, mixed precision networks achieve better performance than networks with homogeneous bitwidth for the same size constraint. Since choosing the optimal bitwidths is not straight forward, training methods, which can learn them, are desirable. Differentiable quantization with straight-through gradients allows to learn the quantizer's parameters using gradient methods. We show that a suited parametrization of the quantizer is the key to achieve a stable training and a good final performance. Specifically, we propose to parametrize the quantizer with the step size and dynamic range. The bitwidth can then be inferred from them. Other parametrizations, which explicitly use the bitwidth, consistently perform worse. We confirm our findings with experiments on CIFAR-10 and ImageNet and we obtain mixed precision DNNs with learned quantization parameters, achieving state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- Overcoming Oscillations in Quantization-Aware TrainingMarkus Nagel, Marios Fournarakis, Yelysei Bondarenko, Tijmen BlankevoortICML 2022 · 被引用 163 次
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel 等NeurIPS 2022 · 被引用 154 次
- Pruning vs Quantization: Which is Better?Andrey Kuzmin, Markus Nagel, Mart van Baalen, Arash Behboodi 等NeurIPS 2023 · 被引用 152 次
- Bayesian Bits: Unifying Quantization and PruningMart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad 等NeurIPS 2020 · 被引用 149 次
它引用的顶会 Paper1
相关 Paper
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 被引用 95 次
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive ResolutionZhaoyang Zhang, Wenqi Shao, Jinwei Gu, Xiaogang Wang 等ICML 2021 · 被引用 36 次
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu 等ICML 2022 · 被引用 49 次
- Not All Bits have Equal Value: Heterogeneous Precisions via Trainable NoisePedro Savarese, Xin Yuan, Yanjing Li, Michael MaireNeurIPS 2022 · 被引用 9 次
