One-Shot Model for Mixed-Precision Quantization
Ivan Koryakovskiy, Alexandra Yakovleva, Valentin Buchnev, Temur Isaev, Gleb Odinokikh
摘要
Neural network quantization is a popular approach for model compression. Modern hardware supports quantization in mixed-precision mode, which allows for greater compression rates but adds the challenging task of searching for the optimal bit width. The majority of existing searchers find a single mixed-precision architecture. To select an architecture that is suitable in terms of performance and resource consumption, one has to restart searching multiple times. We focus on a specific class of methods that find tensor bit width using gradient-based optimization. First, we theoretically derive several methods that were empirically proposed earlier. Second, we present a novel One-Shot method that finds a diverse set of Pareto-front architectures in O(1) time. For large models, the proposed method is 5 times more efficient than existing methods. We verify the method on two classification and super-resolution models and show above 0.93 correlation score between the predicted and actual model performance. The Pareto-front architecture selection is straightforward and takes only 20 to 40 supernet evaluations, which is the new state-of-the-art result to the best of our knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data FormatChao Fang, Man Shi, Robin Geens, Arne Symons 等HPCA 2025 · 被引用 15 次
- PikeLPN: Mitigating Overlooked Inefficiencies of Low-Precision Neural NetworksMarina Neseem, Conor McCullough, Randy Hsin, Chas Leichner 等CVPR 2024 · 被引用 3 次
- Efficient Mixed Precision Quantization in Graph Neural NetworksSamir Moustafa, Nils M. Kriege, Wilfried N. GanstererICDE 2025 · 被引用 2 次
它引用的顶会 Paper14
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang 等ICLR 2021 · 被引用 619 次
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami 等NeurIPS 2020 · 被引用 434 次
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- AutoQ: Automated Kernel-Wise Neural Network QuantizationQian Lou, Feng Guo, Minje Kim, Lantao Liu 等ICLR 2020 · 被引用 121 次
相关 Paper
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie 等CVPR 2024 · 被引用 5 次
- OMPQ: Orthogonal Mixed Precision QuantizationYuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang 等AAAI 2023 · 被引用 56 次
- Towards Mixed-Precision Quantization of Neural Networks via Constrained OptimizationWeihan Chen, Peisong Wang, Jian ChengICCV 2021 · 被引用 91 次
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 被引用 95 次
