Distillation-Based Training for Multi-Exit Architectures
Mary Phuong, Christoph Lampert
Abstract
Multi-exit architectures, in which a stack of processing layers is interleaved with early output layers, allow the processing of a test example to stop early and thus save computation time and/or energy. In this work, we propose a new training procedure for multi-exit architectures based on the principle of knowledge distillation. The method encourages early exits to mimic later, more accurate exits, by matching their probability outputs. Experiments on CIFAR100 and ImageNet show that distillation-based training significantly improves the accuracy of early exits while maintaining state-of-the-art accuracy for late ones. The method is particularly beneficial when training data is limited and also allows a straight-forward extension to semi-supervised learning, i.e. make use also of unlabeled data at training time. Moreover, it takes only a few lines to implement and imposes almost no computational overhead at training time, and none at all at test time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a0b6fc5-228c-475e-94e5-9f8ea8311af3Cited by top-tier papers30
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang et al.NeurIPS 2023 · 205 citations
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li et al.CVPR 2022 · 88 citations
- Zero Time Waste: Recycling Predictions in Early Exit Neural NetworksMaciej Wolczyk, Bartosz Wójcik, Klaudia Balazy, Igor T. Podolak et al.NeurIPS 2021 · 78 citations
- SplitGP: Achieving Both Generalization and Personalization in Federated LearningDong-Jun Han, Do-Yeon Kim, Minseok Choi, Christopher G. Brinton et al.INFOCOM 2023 · 43 citations
- Self-Regulation for Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.ICCV 2021 · 42 citations
Builds on1
Related papers
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 26 citations
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 11 citations
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu et al.AAAI 2023 · 40 citations
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 21 citations
- DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous DevicesLehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu et al.KDD 2025
