Searching for Low-Bit Weights in Quantized Neural Networks
Zhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu, Chao Xu, Dacheng Tao, Chang Xu
摘要
Quantized neural networks with low-bit weights and activations are attractive for developing AI accelerators. However, the quantization functions used in most conventional quantization methods are non-differentiable, which increases the optimization difficulty of quantized networks. Compared with full-precision parameters (i.e., 32-bit floating numbers), low-bit values are selected from a much smaller set. For example, there are only 16 possibilities in 4-bit space. Thus, we present to regard the discrete weights in an arbitrary quantized neural network as searchable variables, and utilize a differential method to search them accurately. In particular, each weight is represented as a probability distribution over the discrete value set. The probabilities are optimized during training and the values with the highest probability are selected to establish the desired quantized network. Experimental results on benchmarks demonstrate that the proposed method is able to produce quantized neural networks with higher performance over the state-of-the-art methods on both image classification and super-resolution tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- SCOP: Scientific Control for Reliable Neural Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Dacheng Tao 等NeurIPS 2020 · 被引用 208 次
- ReCU: Reviving the Dead Weights in Binary Neural NetworksZihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen 等ICCV 2021 · 被引用 102 次
- Learning Frequency Domain Approximation for Binary Neural NetworksYixing Xu, Kai Han, Chang Xu, Yehui Tang 等NeurIPS 2021 · 被引用 64 次
- Training Stronger Baselines for Learning to OptimizeTianlong Chen, Weiyi Zhang, Jingyang Zhou, Shiyu Chang 等NeurIPS 2020 · 被引用 61 次
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante 等ICLR 2022 · 被引用 57 次
它引用的顶会 Paper4
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- Bayesian Optimized 1-Bit CNNsJiaxin Gu, Junhe Zhao, Xiaolong Jiang, Baochang Zhang 等ICCV 2019 · 被引用 57 次
- Proximal Mean-Field for Neural Network QuantizationThalaiyasingam Ajanthan, Puneet K. Dokania, Richard Hartley, Philip H. S. TorrICCV 2019 · 被引用 21 次
- Cogradient Descent for Bilinear OptimizationLi'an Zhuo, Baochang Zhang, Linlin Yang, Hanlin Chen 等CVPR 2020
相关 Paper
- Learnable Lookup Table for Neural Network QuantizationLongguang Wang, Xiaoyu Dong, Yingqian Wang, Li Liu 等CVPR 2022 · 被引用 52 次
- BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network QuantizationHuanrui Yang, Lin Duan, Yiran Chen, Hai LiICLR 2021 · 被引用 83 次
- Instance-Aware Dynamic Neural Network QuantizationZhenhua Liu, Yunhe Wang, Kai Han, Siwei Ma 等CVPR 2022 · 被引用 38 次
- Rethinking Differentiable Search for Mixed-Precision Neural NetworksZhaowei Cai, Nuno VasconcelosCVPR 2020
- SDQ: Stochastic Differentiable Quantization with Mixed PrecisionXijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu 等ICML 2022 · 被引用 49 次
