Harmonized Dense Knowledge Distillation Training for Multi-Exit Architectures
Xinglu Wang, Yingming Li
Abstract
Multi-exit architectures, in which a sequence of intermediate classifiers are introduced at different depths of the feature layers, perform adaptive computation by early exiting ``easy" samples to speed up the inference. In this paper, a novel Harmonized Dense Knowledge Distillation (HDKD) training method for multi-exit architecture is designed to encourage each exit to flexibly learn from all its later exits. In particular, a general dense knowledge distillation training objective is proposed to incorporate all possible beneficial supervision information for multi-exit learning, where a harmonized weighting scheme is designed for the multi-objective optimization problem consisting of multi-exit classification loss and dense distillation loss. A bilevel optimization algorithm is introduced for alternatively updating the weights of multiple objectives and the multi-exit network parameters. Specifically, the loss weighting parameters are optimized with respect to its performance on validation set by gradient descent. Experiments on CIFAR100 and ImageNet show that the HDKD strategy harmoniously improves the performance of the state-of-the-art multi-exit neural networks. Moreover, this method does not require within architecture modifications and can be effectively combined with other previously-proposed training techniques and further boosts the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd26d92b-9c66-4150-8acc-738317de596fCited by top-tier papers4
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li et al.CVPR 2022 · 88 citations
- Auditing Membership Leakages of Multi-Exit NetworksZheng Li, Yiyong Liu, Xinlei He, Ning Yu et al.CCS 2022 · 19 citations
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 11 citations
- Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GameKeyizhi Xu, Chi Zhang, Zhan Chen, Zhongyuan Wang et al.CVPR 2025
Builds on4
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 205 citations
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang et al.ICCV 2019 · 152 citations
Related papers
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous DevicesLehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu et al.KDD 2025
- Multi-Knowledge Aggregation and Transfer for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2022 · 11 citations
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
- How to Train Your Multi-Exit Model? Analyzing the Impact of Training StrategiesPiotr Kubaty, Bartosz Wójcik, Bartlomiej Krzepkowski, Monika Michaluk et al.ICML 2025
