AutoQ: Automated Kernel-Wise Neural Network Quantization
Qian Lou, Feng Guo, Minje Kim, Lantao Liu, Lei Jiang
Abstract
Network quantization is one of the most hardware friendly techniques to enable the deployment of convolutional neural networks (CNNs) on low-power mobile devices. Recent network quantization techniques quantize each weight kernel in a convolutional layer independently for higher inference accuracy, since the weight kernels in a layer exhibit different variances and hence have different amounts of redundancy. The quantization bitwidth or bit number (QBN) directly decides the inference accuracy, latency, energy and hardware overhead. To effectively reduce the redundancy and accelerate CNN inferences, various weight kernels should be quantized with different QBNs. However, prior works use only one QBN to quantize each convolutional layer or the entire CNN, because the design space of searching a QBN for each weight kernel is too large. The hand-crafted heuristic of the kernel-wise QBN search is so sophisticated that domain experts can obtain only sub-optimal results. It is difficult for even deep reinforcement learning (DRL) Deep Deterministic Policy Gradient (DDPG)-based agents to find a kernel-wise QBN configuration that can achieve reasonable inference accuracy. In this paper, we propose a hierarchical-DRL-based kernel-wise network quantization technique, AutoQ, to automatically search a QBN for each weight kernel, and choose another QBN for each activation layer. Compared to the models quantized by the state-of-the-art DRL-based schemes, on average, the same models quantized by AutoQ reduce the inference latency by 54.06%, and decrease the inference energy consumption by 50.69%, while achieving the same inference accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85624d69-e1bf-4a6d-b377-273f0e5fc37cCited by top-tier papers23
- Towards Mixed-Precision Quantization of Neural Networks via Constrained OptimizationWeihan Chen, Peisong Wang, Jian ChengICCV 2021 · 91 citations
- FlexRound: Learnable Rounding based on Element-wise Division for Post-Training QuantizationJung Hyun Lee, Jeonghoon Kim, Se Jung Kwon, Dongsoo LeeICML 2023 · 56 citations
- EMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationPeijie Dong, Lujun Li, Zimian Wei, Xin Niu et al.ICCV 2023 · 51 citations
- AutoPrivacy: Automated Layer-wise Parameter Selection for Secure Neural Network InferenceQian Lou, Song Bian, Lei JiangNeurIPS 2020 · 41 citations
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma et al.CVPR 2022 · 38 citations
Related papers
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- RaQu: An automatic high-utilization CNN quantization and mapping framework for general-purpose RRAM AcceleratorSongyun Qu, Bing Li, Ying Wang, Dawen Xu et al.DAC 2020 · 40 citations
- PowerQuant: Automorphism Search for Non-Uniform QuantizationEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyICLR 2023 · 1 citation
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
- Simple Augmentation Goes a Long Way: ADRL for DNN QuantizationLin Ning, Guoyang Chen, Weifeng Zhang, Xipeng ShenICLR 2021 · 1 citation
