Kernel Based Progressive Distillation for Adder Neural Networks
Yixing Xu, Chang Xu, Xinghao Chen, Wei Zhang, Chunjing Xu, Yunhe Wang
Abstract
Adder Neural Networks (ANNs) which only contain additions bring us a new way of developing deep neural networks with low energy consumption. Unfortunately, there is an accuracy drop when replacing all convolution filters by adder filters. The main reason here is the optimization difficulty of ANNs using -norm, in which the estimation of gradient in back propagation is inaccurate. In this paper, we present a novel method for further improving the performance of ANNs without increasing the trainable parameters via a progressive kernel based knowledge distillation (PKKD) method. A convolutional neural network (CNN) with the same architecture is simultaneously initialized and trained as a teacher network, features and weights of ANN and CNN will be transformed to a new space to eliminate the accuracy drop. The similarity is conducted in a higher-dimensional space to disentangle the difference of their distributions using a kernel based method. Finally, the desired ANN is learned based on the information from both the ground-truth and teacher, progressively. The effectiveness of the proposed method for learning ANN with higher performance is then well-verified on several benchmarks. For instance, the ANN-50 trained using the proposed PKKD method obtains a 76.8% top-1 accuracy on ImageNet dataset, which is 0.6% higher than that of the ResNet-50.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang et al.NeurIPS 2021 · 71 citations
- Learning Frequency Domain Approximation for Binary Neural NetworksYixing Xu, Kai Han, Chang Xu, Yehui Tang et al.NeurIPS 2021 · 64 citations
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist NeuronsBohang Zhang, Tianle Cai, Zhou Lu, Di He et al.ICML 2021 · 62 citations
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou et al.NeurIPS 2024 · 47 citations
Builds on7
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- AutoGAN: Neural Architecture Search for Generative Adversarial NetworksXinyu Gong, Shiyu Chang, Yifan Jiang, Zhangyang WangICCV 2019 · 286 citations
- Knowledge Distillation via Route Constrained OptimizationXiao Jin, Baoyun Peng, Yichao Wu, Yu Liu et al.ICCV 2019 · 196 citations
- Efficient Residual Dense Block Search for Image Super-ResolutionDehua Song, Chang Xu, Xu Jia, Yiyi Chen et al.AAAI 2020 · 145 citations
Related papers
- Handling Long-tailed Feature Distribution in AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 3 citations
- AdderNet: Do We Really Need Multiplications in Deep Learning?Hanting Chen, Yunhe Wang, Chunjing Xu, Boxin Shi et al.CVPR 2020
- Towards Stable and Robust AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 11 citations
- Adder Attention for Vision TransformerHan Shu, Jiahao Wang, Hanting Chen, Lin Li et al.NeurIPS 2021 · 23 citations
- AdderSR: Towards Energy Efficient Image Super-ResolutionDehua Song, Yunhe Wang, Hanting Chen, Chang Xu et al.CVPR 2021
