Harmonized Dense Knowledge Distillation Training for Multi-Exit Architectures
Xinglu Wang, Yingming Li
摘要
Multi-exit architectures, in which a sequence of intermediate classifiers are introduced at different depths of the feature layers, perform adaptive computation by early exiting ``easy" samples to speed up the inference. In this paper, a novel Harmonized Dense Knowledge Distillation (HDKD) training method for multi-exit architecture is designed to encourage each exit to flexibly learn from all its later exits. In particular, a general dense knowledge distillation training objective is proposed to incorporate all possible beneficial supervision information for multi-exit learning, where a harmonized weighting scheme is designed for the multi-objective optimization problem consisting of multi-exit classification loss and dense distillation loss. A bilevel optimization algorithm is introduced for alternatively updating the weights of multiple objectives and the multi-exit network parameters. Specifically, the loss weighting parameters are optimized with respect to its performance on validation set by gradient descent. Experiments on CIFAR100 and ImageNet show that the HDKD strategy harmoniously improves the performance of the state-of-the-art multi-exit neural networks. Moreover, this method does not require within architecture modifications and can be effectively combined with other previously-proposed training techniques and further boosts the performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li 等CVPR 2022 · 被引用 88 次
- Auditing Membership Leakages of Multi-Exit NetworksZheng Li, Yiyong Liu, Xinlei He, Ning Yu 等CCS 2022 · 被引用 19 次
- NEO-KD: Knowledge-Distillation-Based Adversarial Training for Robust Multi-Exit Neural NetworksSeokil Ham, Jungwuk Park, Dong-Jun Han, Jaekyun MoonNeurIPS 2023 · 被引用 11 次
- Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GameKeyizhi Xu, Chi Zhang, Zhan Chen, Zhongyuan Wang 等CVPR 2025
它引用的顶会 Paper4
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
相关 Paper
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous DevicesLehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu 等KDD 2025
- Multi-Knowledge Aggregation and Transfer for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2022 · 被引用 11 次
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li 等EMNLP 2023 · 被引用 5 次
- How to Train Your Multi-Exit Model? Analyzing the Impact of Training StrategiesPiotr Kubaty, Bartosz Wójcik, Bartlomiej Krzepkowski, Monika Michaluk 等ICML 2025
