Distillation-Based Training for Multi-Exit Architectures
Mary Phuong, Christoph Lampert
摘要
Multi-exit architectures, in which a stack of processing layers is interleaved with early output layers, allow the processing of a test example to stop early and thus save computation time and/or energy. In this work, we propose a new training procedure for multi-exit architectures based on the principle of knowledge distillation. The method encourages early exits to mimic later, more accurate exits, by matching their probability outputs. Experiments on CIFAR100 and ImageNet show that distillation-based training significantly improves the accuracy of early exits while maintaining state-of-the-art accuracy for late ones. The method is particularly beneficial when training data is limited and also allows a straight-forward extension to semi-supervised learning, i.e. make use also of unlabeled data at training time. Moreover, it takes only a few lines to implement and imposes almost no computational overhead at training time, and none at all at test time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang 等NeurIPS 2023 · 被引用 205 次
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li 等CVPR 2022 · 被引用 88 次
- Zero Time Waste: Recycling Predictions in Early Exit Neural NetworksMaciej Wolczyk, Bartosz Wójcik, Klaudia Balazy, Igor T. Podolak 等NeurIPS 2021 · 被引用 78 次
- SplitGP: Achieving Both Generalization and Personalization in Federated LearningDong-Jun Han, Do-Yeon Kim, Minseok Choi, Christopher G. Brinton 等INFOCOM 2023 · 被引用 43 次
- Self-Regulation for Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua 等ICCV 2021 · 被引用 42 次
它引用的顶会 Paper1
相关 Paper
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 被引用 26 次
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 被引用 11 次
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu 等AAAI 2023 · 被引用 40 次
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 被引用 21 次
- DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous DevicesLehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu 等KDD 2025
