Improved Techniques for Training Adaptive Deep Networks
Hao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang, Gao Huang
摘要
Adaptive inference is a promising technique to improve the computational efficiency of deep models at test time. In contrast to static models which use the same computation graph for all instances, adaptive networks can dynamically adjust their structure conditioned on each input. While existing research on adaptive inference mainly focuses on designing more advanced architectures, this paper investigates how to train such networks more effectively. Specifically, we consider a typical adaptive deep network with multiple intermediate classifiers. We present three techniques to improve its training efficacy from two aspects: 1) a Gradient Equilibrium algorithm to resolve the conflict of learning of different classifiers; 2) an Inline Subnetwork Collaboration approach and a One-for-all Knowledge Distillation algorithm to enhance the collaboration among classifiers. On multiple datasets (CIFAR-10, CIFAR-100 and ImageNet), we show that the proposed approach consistently leads to further improved efficiency on top of state-of-the-art adaptive deep networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang 等NeurIPS 2021 · 被引用 283 次
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song 等NeurIPS 2020 · 被引用 179 次
- Spatially-Adaptive Image Restoration using Distortion-Guided NetworksKuldeep Purohit, Maitreya Suin, A. N. Rajagopalan, Vishnu Naresh BoddetiICCV 2021 · 被引用 156 次
- Zero Time Waste: Recycling Predictions in Early Exit Neural NetworksMaciej Wolczyk, Bartosz Wójcik, Klaudia Balazy, Igor T. Podolak 等NeurIPS 2021 · 被引用 78 次
相关 Paper
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 被引用 21 次
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 被引用 5 次
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 被引用 26 次
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu 等CVPR 2020
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen 等AAAI 2023 · 被引用 16 次
