AdderNet: Do We Really Need Multiplications in Deep Learning?
Hanting Chen, Yunhe Wang, Chunjing Xu, Boxin Shi, Chao Xu, Qi Tian, Chang Xu
Abstract
Compared with cheap addition operation, multiplication operation is of much higher computation complexity. The widely-used convolutions in deep neural networks are exactly cross-correlation to measure the similarity between input feature and convolution filters, which involves massive multiplications between float values. In this paper, we present adder networks (AdderNets) to trade these massive multiplications in deep neural networks, especially convolutional neural networks (CNNs), for much cheaper additions to reduce computation costs. In AdderNets, we take the 1 -norm distance between filters and input feature as the output response. The influence of this new similarity measure on the optimization of neural network have been thoroughly analyzed. To achieve a better performance, we develop a special back-propagation approach for AdderNets by investigating the full-precision gradient. We then propose an adaptive learning rate strategy to enhance the training procedure of AdderNets according to the magnitude of each neuron's gradient. As a result, the proposed AdderNets can achieve 74.9% Top-1 accuracy 91.7% Top-5 accuracy using ResNet-50 on the Im-ageNet dataset without any multiplication in convolution layer. The codes are publicly available at: https:// github.com/huaweinoah/AdderNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa9c238a-1bbc-4135-91da-a5376e783d6eCited by top-tier papers39
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
- TopFormer: Token Pyramid Transformer for Mobile Semantic SegmentationWenqiang Zhang, Zilong Huang, Guozhong Luo, Tao Chen et al.CVPR 2022 · 313 citations
- Differentiable Spike: Rethinking Gradient-Descent for Training Spiking Neural NetworksYuhang Li, Yufei Guo, Shanghang Zhang, Shikuang Deng et al.NeurIPS 2021 · 288 citations
- Rotated Binary Neural NetworkMingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang et al.NeurIPS 2020 · 161 citations
- MicroNet: Improving Image Recognition with Extremely Low FLOPsYunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen et al.ICCV 2021 · 108 citations
Related papers
- Winograd Algorithm for AdderNetWenshuo Li, Hanting Chen, Mingqiang Huang, Xinghao Chen et al.ICML 2021 · 8 citations
- Towards Stable and Robust AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 11 citations
- AdderSR: Towards Energy Efficient Image Super-ResolutionDehua Song, Yunhe Wang, Hanting Chen, Chang Xu et al.CVPR 2021
- Kernel Based Progressive Distillation for Adder Neural NetworksYixing Xu, Chang Xu, Xinghao Chen, Wei Zhang et al.NeurIPS 2020 · 48 citations
- Redistribution of Weights and Activations for AdderNet QuantizationYing Nie, Kai Han, Haikang Diao, Chuanjian Liu et al.NeurIPS 2022 · 14 citations
