Kernel Based Progressive Distillation for Adder Neural Networks
Yixing Xu, Chang Xu, Xinghao Chen, Wei Zhang, Chunjing Xu, Yunhe Wang
摘要
Adder Neural Networks (ANNs) which only contain additions bring us a new way of developing deep neural networks with low energy consumption. Unfortunately, there is an accuracy drop when replacing all convolution filters by adder filters. The main reason here is the optimization difficulty of ANNs using -norm, in which the estimation of gradient in back propagation is inaccurate. In this paper, we present a novel method for further improving the performance of ANNs without increasing the trainable parameters via a progressive kernel based knowledge distillation (PKKD) method. A convolutional neural network (CNN) with the same architecture is simultaneously initialized and trained as a teacher network, features and weights of ANN and CNN will be transformed to a new space to eliminate the accuracy drop. The similarity is conducted in a higher-dimensional space to disentangle the difference of their distributions using a kernel based method. Finally, the desired ANN is learned based on the information from both the ground-truth and teacher, progressively. The effectiveness of the proposed method for learning ANN with higher performance is then well-verified on several benchmarks. For instance, the ANN-50 trained using the proposed PKKD method obtains a 76.8% top-1 accuracy on ImageNet dataset, which is 0.6% higher than that of the ResNet-50.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li 等NeurIPS 2020 · 被引用 99 次
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang 等NeurIPS 2021 · 被引用 71 次
- Learning Frequency Domain Approximation for Binary Neural NetworksYixing Xu, Kai Han, Chang Xu, Yehui Tang 等NeurIPS 2021 · 被引用 64 次
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist NeuronsBohang Zhang, Tianle Cai, Zhou Lu, Di He 等ICML 2021 · 被引用 62 次
- ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less ReparameterizationHaoran You, Yipin Guo, Yichao Fu, Wei Zhou 等NeurIPS 2024 · 被引用 47 次
它引用的顶会 Paper7
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang 等ICCV 2019 · 被引用 427 次
- AutoGAN: Neural Architecture Search for Generative Adversarial NetworksXinyu Gong, Shiyu Chang, Yifan Jiang, Zhangyang WangICCV 2019 · 被引用 286 次
- Knowledge Distillation via Route Constrained OptimizationXiao Jin, Baoyun Peng, Yichao Wu, Yu Liu 等ICCV 2019 · 被引用 196 次
- Efficient Residual Dense Block Search for Image Super-ResolutionDehua Song, Chang Xu, Xu Jia, Yiyi Chen 等AAAI 2020 · 被引用 145 次
相关 Paper
- Handling Long-tailed Feature Distribution in AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 被引用 3 次
- AdderNet: Do We Really Need Multiplications in Deep Learning?Hanting Chen, Yunhe Wang, Chunjing Xu, Boxin Shi 等CVPR 2020
- Towards Stable and Robust AdderNetsMinjing Dong, Yunhe Wang, Xinghao Chen, Chang XuNeurIPS 2021 · 被引用 11 次
- Adder Attention for Vision TransformerHan Shu, Jiahao Wang, Hanting Chen, Lin Li 等NeurIPS 2021 · 被引用 23 次
- AdderSR: Towards Energy Efficient Image Super-ResolutionDehua Song, Yunhe Wang, Hanting Chen, Chang Xu 等CVPR 2021
