CSQ: Growing Mixed-Precision Quantization Scheme with Bi-level Continuous Sparsification
Lirui Xiao, Huanrui Yang, Zhen Dong, Kurt Keutzer, Li Du, Shanghang Zhang
摘要
Mixed-precision quantization has been widely applied on deep neural networks (DNNs) as it leads to significantly better efficiency-accuracy tradeoffs compared to uniform quantization. Meanwhile, determining the exact precision of each layer remains challenging. Previous attempts on bit-level regularization and pruning-based dynamic precision adjustment during training suffer from noisy gradients and unstable convergence. In this work, we propose Continuous Sparsification Quantization (CSQ), a bit-level training method to search for mixed-precision quantization schemes with improved stability. CSQ stabilizes the bit-level mixed-precision training process with a bi-level gradual continuous sparsification on both the bit values of the quantized weights and the bit selection in determining the quantization precision of each layer. The continuous sparsification scheme enables fully-differentiable training without gradient approximation while achieving an exact quantized model in the end. A budget-aware regularization of total model size enables the dynamic growth and pruning of each layer's precision towards a mixed-precision quantization scheme of the desired size. Extensive experiments show CSQ achieves better efficiencyaccuracy tradeoff than previous methods on multiple models and datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- QD-BEV : Quantization-aware View-guided Distillation for Multi-view 3D Object DetectionYifan Zhang, Zhen Dong, Huanrui Yang, Ming Lu 等ICCV 2023 · 被引用 15 次
- Lossy and Lossless (L2) Post-training Model Size CompressionYumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong 等ICCV 2023 · 被引用 5 次
- MSQ: Memory-Efficient Bit Sparsification QuantizationSeokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang 等ICCV 2025 · 被引用 2 次
- NoisyQuant: Noisy Bias-Enhanced Post-Training Activation Quantization for Vision TransformersYijiang Liu, Huanrui Yang, Zhen Dong, Kurt Keutzer 等CVPR 2023
它引用的顶会 Paper8
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney 等ICCV 2019 · 被引用 645 次
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami 等NeurIPS 2020 · 被引用 434 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 被引用 162 次
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
相关 Paper
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 被引用 95 次
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu 等ICML 2022 · 被引用 49 次
- Any-Precision Deep Neural NetworksHaichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang 等AAAI 2021 · 被引用 79 次
- Cluster-Promoting Quantization with Bit-Drop for Minimizing Network Quantization LossJung Hyun Lee, Jihun Yun, Sung Ju Hwang, Eunho YangICCV 2021 · 被引用 17 次
