Residual Distillation: Towards Portable Deep Neural Networks without Shortcuts
Guilin Li, Junlei Zhang, Yunhe Wang, Chuanjian Liu, Matthias H. Y. Tan, Yunfeng Lin, Wei Zhang, Jiashi Feng, Tong Zhang
Abstract
By transferring both features and gradients between different layers, shortcut connections explored by ResNets allow us to effectively train very deep neural networks up to hundreds of layers. However, the additional computation costs induced by those shortcuts are often overlooked. For example, during online inference, the shortcuts in ResNet-50 account for about 40 percent of the entire memory usage on feature maps, because the features in the preceding layers cannot be released until the subsequent calculation is completed. In this work, for the first time, we consider training the CNN models with shortcuts and deploying them without. In particular, we propose a novel joint-training framework to train plain CNN by leveraging the gradients of the ResNet counterpart. During forward step, the feature maps of the early stages of plain CNN are passed through later stages of both itself and the ResNet counterpart to calculate the loss. During backpropagation, gradients calculated from a mixture of these two parts are used to update the plainCNN network to solve the gradient vanishing problem. Extensive experiments on ImageNet/CIFAR10/CIFAR100 demonstrate that the plainCNN network without shortcuts generated by our approach can achieve the same level of accuracy as that of the ResNet baseline while achieving about 1.4× speed-up and 1.25× memory reduction. We also verified the feature transferability of our ImageNet pretrained plain-CNN network by fine-tuning it on MIT 67 and Caltech 101. Our results show that the performance of the plain-CNN is slightly higher than that of its baseline ResNet-50 on these two datasets. The code will be available at https://github.com/leoozy/JointRD_Neurips2020 and the MindSpore code will be available at https://www.mindspore.cn/resources/hub. © 2020 Neural information processing systems foundation. All rights reserved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6ee89f5-230e-4a4a-9b6d-71463420fabcCited by top-tier papers6
- VanillaNet: the Power of Minimalism in Deep LearningHanting Chen, Yunhe Wang, Jianyuan Guo, Dacheng TaoNeurIPS 2023 · 228 citations
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li et al.CVPR 2024 · 93 citations
- Undistillable: Making A Nasty Teacher That CANNOT teach studentsHaoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You et al.ICLR 2021 · 58 citations
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 20 citations
- Instance Smoothed Contrastive Learning for Unsupervised Sentence EmbeddingHongliang He, Junlei Zhang, Zhenzhong Lan, Yue ZhangAAAI 2023 · 10 citations
Builds on9
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- AutoGAN: Neural Architecture Search for Generative Adversarial NetworksXinyu Gong, Shiyu Chang, Yifan Jiang, Zhangyang WangICCV 2019 · 286 citations
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang et al.ICML 2020 · 122 citations
- Co-Evolutionary Compression for Unpaired Image TranslationHan Shu, Yunhe Wang, Xu Jia, Kai Han et al.ICCV 2019 · 93 citations
Related papers
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski et al.ICLR 2026
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng et al.VLDB 2022 · 39 citations
- Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural NetworksYufei Guo, Yuanpei Chen, Zecheng Hao, Weihang Peng et al.NeurIPS 2024 · 23 citations
- Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural NetworksWoochul Kang, Daeyeon KimNeurIPS 2021 · 5 citations
- Prior Gradient Mask Guided Pruning-Aware Fine-TuningLinhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan et al.AAAI 2022 · 44 citations
