SDQ: Stochastic Differentiable Quantization with Mixed Precision
Xijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu, Xianghong Hu, Jeffry Wicaksana, Eric P. Xing, Kwang-Ting Cheng
摘要
In order to deploy deep models in a computationally efficient manner, model quantization approaches have been frequently used. In addition, as new hardware that supports mixed bitwidth arithmetic operations, recent research on mixed precision quantization (MPQ) begins to fully leverage the capacity of representation by searching optimized bitwidths for different layers and modules in a network. However, previous studies mainly search the MPQ strategy in a costly scheme using reinforcement learning, neural architecture search, etc., or simply utilize partial prior knowledge for bitwidth assignment, which might be biased on locality of information and is sub-optimal. In this work, we present a novel Stochastic Differentiable Quantization (SDQ) method that can automatically learn the MPQ strategy in a more flexible and globallyoptimized space with smoother gradient approximation. Particularly, Differentiable Bitwidth Parameters (DBPs) are employed as the probability factors in stochastic quantization between adjacent bitwidth choices. After the optimal MPQ strategy is acquired, we further train our network with Entropy-aware Bin Regularization and knowledge distillation. We extensively evaluate our method for several networks on different hardware (GPUs and FPGA) and datasets. SDQ outperforms all state-of-the-art mixed or single precision quantization with a lower bitwidth and is even better than the full-precision counterparts across various ResNet and MobileNet families, demonstrating its effectiveness and superiority. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- PTQD: Accurate Post-Training Quantization for Diffusion ModelsYefei He, Luping Liu, Jing Liu, Weijia Wu 等NeurIPS 2023 · 被引用 219 次
- EQ-Net: Elastic Quantization Neural NetworksKe Xu, Lei Han, Ye Tian, Shangshang Yang 等ICCV 2023 · 被引用 21 次
- Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersPingcheng Dong, Yonghao Tan, Dong Zhang, Tianwei Ni 等DAC 2024 · 被引用 12 次
- SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin RegularizationChen Tang, Kai Ouyang, Zenghao Chai, Yunpeng Bai 等ACM MM 2023 · 被引用 11 次
- MetaMix: Meta-State Precision Searcher for Mixed-Precision Activation QuantizationHan-Byul Kim, Joo Hyung Lee, Sungjoo Yoo, Hong-Seok KimAAAI 2024 · 被引用 10 次
它引用的顶会 Paper11
- Additive Powers-of-Two Quantization: An Efficient Non-uniform Discretization for Neural NetworksYuhang Li, Xin Dong, Wei WangICLR 2020 · 被引用 315 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- Robust Quantization: One Model to Rule Them AllMoran Shkolnik, Brian Chmiel, Ron Banner, Gil Shomron 等NeurIPS 2020 · 被引用 103 次
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 被引用 95 次
相关 Paper
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive ResolutionZhaoyang Zhang, Wenqi Shao, Jinwei Gu, Xiaogang Wang 等ICML 2021 · 被引用 36 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- MSQ: Memory-Efficient Bit Sparsification QuantizationSeokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang 等ICCV 2025 · 被引用 2 次
- Rethinking Differentiable Search for Mixed-Precision Neural NetworksZhaowei Cai, Nuno VasconcelosCVPR 2020
