Instance-Aware Dynamic Neural Network Quantization
Zhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma, Wen Gao
Abstract
Quantization is an effective way to reduce the memory and computational costs of deep neural networks in which the full-precision weights and activations are represented using low-bit values. The bit-width for each layer in most of existing quantization methods is static, i.e., the same for all samples in the given dataset. However, natural images are of huge diversity with abundant content and using such a universal quantization configuration for all samples is not an optimal strategy. In this paper, we present to conduct the low-bit quantization for each image individually, and develop a dynamic quantization scheme for exploring their optimal bit-widths. To this end, a lightweight bit-controller is established and trained jointly with the given neural network to be quantized. During inference, the quantization configuration for an arbitrary image will be determined by the bit-widths generated by the controller, e.g., an image with simple texture will be allocated with lower bits and computational complexity and vice versa. Experimental results conducted on benchmarks demonstrate the effectiveness of the proposed dynamic quantization method for achieving state-of-art performance in terms of accuracy and computational complexity. The code will be available at https://github.com/huawei-noah/Efficient-Computing and https://gitee.com/mindspore/models/tree/master/research/cv/DynamicQuant.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f3c8df9-dede-486c-9be7-32b5304b1213Cited by top-tier papers11
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim et al.NeurIPS 2023 · 109 citations
- Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt TuningKun Ding, Haojian Zhang, Qiang Yu, Ying Wang et al.AAAI 2024 · 8 citations
- Thinking in Granularity: Dynamic Quantization for Image Super-Resolution by Intriguing Multi-Granularity CluesMingshen Wang, Zhao Zhang, Feng Li, Ke Xu et al.AAAI 2025 · 4 citations
- DQT: Dynamic Quantization Training via Dequantization-Free Nested Integer ArithmeticHazem Hesham Yousef Shalby, Fabrizio Pittorino, Francesca Palermo, Diana Trojaniello et al.AAAI 2026 · 2 citations
- Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsHongjae Lee, Myungjun Son, Dongjea Kang, Seung-Won JungICCV 2025 · 2 citations
Builds on11
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami et al.NeurIPS 2020 · 434 citations
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama et al.ICLR 2020 · 159 citations
Related papers
- Dynamic Network Quantization for Efficient Video InferenceXimeng Sun, Rameswar Panda, Chun-Fu (Richard) Chen, Aude Oliva et al.ICCV 2021 · 56 citations
- Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference ApproachChen Tang, Haoyu Zhai, Kai Ouyang, Zhi Wang et al.ACM MM 2022 · 15 citations
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
- Distribution-Aware Adaptive Multi-Bit QuantizationSijie Zhao, Tao Yue, Xuemei HuCVPR 2021
- Differentiable Dynamic Quantization with Mixed Precision and Adaptive ResolutionZhaoyang Zhang, Wenqi Shao, Jinwei Gu, Xiaogang Wang et al.ICML 2021 · 36 citations
