FracBits: Mixed Precision Quantization via Fractional Bit-Widths
Linjie Yang, Qing Jin
摘要
Model quantization helps to reduce model size and latency of deep neural networks. Mixed precision quantization is favorable with customized hardwares supporting arithmetic operations at multiple bit-widths to achieve maximum efficiency. We propose a novel learning-based algorithm to derive mixed precision models end-to-end under target computation constraints and model sizes. During the optimization, the bit-width of each layer / kernel in the model is at a fractional status of two consecutive bit-widths which can be adjusted gradually. With a differentiable regularization term, the resource constraints can be met during the quantization-aware training which results in an optimized mixed precision model. Our final models achieve comparable or better performance than previous quantization methods with mixed precision on MobilenetV1/V2, ResNet18 under different resource constraints on ImageNet dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationYunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu 等CVPR 2022 · 被引用 79 次
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante 等ICLR 2022 · 被引用 57 次
- EMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationPeijie Dong, Lujun Li, Zimian Wei, Xin Niu 等ICCV 2023 · 被引用 51 次
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu 等ICML 2022 · 被引用 49 次
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma 等CVPR 2022 · 被引用 38 次
它引用的顶会 Paper3
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- AtomNAS: Fine-Grained End-to-End Neural Architecture SearchJieru Mei, Yingwei Li, Xiaochen Lian, Xiaojie Jin 等ICLR 2020 · 被引用 110 次
相关 Paper
- Mixed-Precision Quantization for Federated Learning on Resource-Constrained Heterogeneous DevicesHuancheng Chen, Haris VikaloCVPR 2024
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive ResolutionZhaoyang Zhang, Wenqi Shao, Jinwei Gu, Xiaogang Wang 等ICML 2021 · 被引用 36 次
- OMPQ: Orthogonal Mixed Precision QuantizationYuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang 等AAAI 2023 · 被引用 56 次
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
