Residual Distillation: Towards Portable Deep Neural Networks without Shortcuts
Guilin Li, Junlei Zhang, Yunhe Wang, Chuanjian Liu, Matthias H. Y. Tan, Yunfeng Lin, Wei Zhang, Jiashi Feng, Tong Zhang
摘要
By transferring both features and gradients between different layers, shortcut connections explored by ResNets allow us to effectively train very deep neural networks up to hundreds of layers. However, the additional computation costs induced by those shortcuts are often overlooked. For example, during online inference, the shortcuts in ResNet-50 account for about 40 percent of the entire memory usage on feature maps, because the features in the preceding layers cannot be released until the subsequent calculation is completed. In this work, for the first time, we consider training the CNN models with shortcuts and deploying them without. In particular, we propose a novel joint-training framework to train plain CNN by leveraging the gradients of the ResNet counterpart. During forward step, the feature maps of the early stages of plain CNN are passed through later stages of both itself and the ResNet counterpart to calculate the loss. During backpropagation, gradients calculated from a mixture of these two parts are used to update the plainCNN network to solve the gradient vanishing problem. Extensive experiments on ImageNet/CIFAR10/CIFAR100 demonstrate that the plainCNN network without shortcuts generated by our approach can achieve the same level of accuracy as that of the ResNet baseline while achieving about 1.4× speed-up and 1.25× memory reduction. We also verified the feature transferability of our ImageNet pretrained plain-CNN network by fine-tuning it on MIT 67 and Caltech 101. Our results show that the performance of the plain-CNN is slightly higher than that of its baseline ResNet-50 on these two datasets. The code will be available at https://github.com/leoozy/JointRD_Neurips2020 and the MindSpore code will be available at https://www.mindspore.cn/resources/hub. © 2020 Neural information processing systems foundation. All rights reserved.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VanillaNet: the Power of Minimalism in Deep LearningHanting Chen, Yunhe Wang, Jianyuan Guo, Dacheng TaoNeurIPS 2023 · 被引用 228 次
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 等CVPR 2024 · 被引用 93 次
- Undistillable: Making A Nasty Teacher That CANNOT teach studentsHaoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You 等ICLR 2021 · 被引用 58 次
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 被引用 20 次
- Instance Smoothed Contrastive Learning for Unsupervised Sentence EmbeddingHongliang He, Junlei Zhang, Zhenzhong Lan, Yue ZhangAAAI 2023 · 被引用 10 次
它引用的顶会 Paper9
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- AutoGAN: Neural Architecture Search for Generative Adversarial NetworksXinyu Gong, Shiyu Chang, Yifan Jiang, Zhangyang WangICCV 2019 · 被引用 286 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
- Co-Evolutionary Compression for Unpaired Image TranslationHan Shu, Yunhe Wang, Xu Jia, Kai Han 等ICCV 2019 · 被引用 93 次
相关 Paper
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski 等ICLR 2026
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng 等VLDB 2022 · 被引用 39 次
- Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural NetworksYufei Guo, Yuanpei Chen, Zecheng Hao, Weihang Peng 等NeurIPS 2024 · 被引用 23 次
- Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural NetworksWoochul Kang, Daeyeon KimNeurIPS 2021 · 被引用 5 次
- Prior Gradient Mask Guided Pruning-Aware Fine-TuningLinhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan 等AAAI 2022 · 被引用 44 次
