Searching for Low-Bit Weights in Quantized Neural Networks
Zhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu, Chao Xu, Dacheng Tao, Chang Xu
Abstract
Quantized neural networks with low-bit weights and activations are attractive for developing AI accelerators. However, the quantization functions used in most conventional quantization methods are non-differentiable, which increases the optimization difficulty of quantized networks. Compared with full-precision parameters (i.e., 32-bit floating numbers), low-bit values are selected from a much smaller set. For example, there are only 16 possibilities in 4-bit space. Thus, we present to regard the discrete weights in an arbitrary quantized neural network as searchable variables, and utilize a differential method to search them accurately. In particular, each weight is represented as a probability distribution over the discrete value set. The probabilities are optimized during training and the values with the highest probability are selected to establish the desired quantized network. Experimental results on benchmarks demonstrate that the proposed method is able to produce quantized neural networks with higher performance over the state-of-the-art methods on both image classification and super-resolution tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a74852e6-b8f9-4017-9a90-b3cc8a4f89c6Cited by top-tier papers26
- SCOP: Scientific Control for Reliable Neural Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Dacheng Tao et al.NeurIPS 2020 · 208 citations
- ReCU: Reviving the Dead Weights in Binary Neural NetworksZihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen et al.ICCV 2021 · 102 citations
- Learning Frequency Domain Approximation for Binary Neural NetworksYixing Xu, Kai Han, Chang Xu, Yehui Tang et al.NeurIPS 2021 · 64 citations
- Training Stronger Baselines for Learning to OptimizeTianlong Chen, Weiyi Zhang, Jingyang Zhou, Shiyu Chang et al.NeurIPS 2020 · 61 citations
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante et al.ICLR 2022 · 57 citations
Builds on4
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang et al.ICCV 2019 · 57 citations
- Proximal Mean-Field for Neural Network QuantizationThalaiyasingam Ajanthan, Puneet K. Dokania, Richard Hartley, Philip H. S. TorrICCV 2019 · 21 citations
- Cogradient Descent for Bilinear OptimizationLi'an Zhuo, Baochang Zhang, Linlin Yang, Hanlin Chen et al.CVPR 2020
Related papers
- Learnable Lookup Table for Neural Network QuantizationLongguang Wang, Xiaoyu Dong, Yingqian Wang, Li Liu et al.CVPR 2022 · 52 citations
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 83 citations
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma et al.CVPR 2022 · 38 citations
- Rethinking Differentiable Search for Mixed-Precision Neural NetworksZhaowei Cai, Nuno VasconcelosCVPR 2020
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu et al.ICML 2022 · 49 citations
