DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous Devices
Lehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu, Yi Xu, Yongxin Tong
Abstract
Early-exit networks (EENs), which adapt their computational depths based on input samples, are widely adopted to accelerate inference in edge computing applications. The effectiveness of EENs relies on difficulty-aware training, which tailors shallow exits for simple samples and deep exits for complex ones. However, existing difficulty-aware training schemes assume centralized environments with sufficient data, which become invalid with real-world edge devices. In this paper, we explore difficulty-aware training in a federated manner, where EENs are collaboratively trained on heterogeneous devices. We observe the cross-model exit unalignment phenomenon, a unique problem when aggregating local EENs into a cohesive global model. To address this problem, we design a novel Difficulty-Aligned Reverse Knowledge Distillation scheme named DarkDistill that preserves the difficulty-specific specialization for aggregating heterogeneous local models. Instead of direct parameter averaging, it trains difficulty-conditional data generators, and selectively transfers generated knowledge of specific difficulty among matched exits of heterogeneous EENs. Evaluations show that Dark-Distill outperforms the state-of-the-arts in both full-parameter and parameter-efficient fine-tuning of EENs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 307d2213-c980-48ac-8159-fd10be5338a3Cited by top-tier papers2
- SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language ModelsYunlong Chu, Minglai Shao, Yuhang Liu, Bing Hao et al.KDD 2026 · 2 citations
- Towards Asynchronous Client Collaboration in Personalized Federated LearningBoyi Liu, Zimu Zhou, Pengfei Gao, Shuo Kang et al.INFOCOM 2026
Builds on25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 957 citations
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningLin Zhang, Li Shen, Liang Ding, Dacheng Tao et al.CVPR 2022 · 339 citations
Related papers
- ScaleFL: Resource-Adaptive Federated Learning with Heterogeneous ClientsFatih Ilhan, Gong Su, Ling LiuCVPR 2023
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 205 citations
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 26 citations
- Unlocking the Non-deterministic Computing Power with Memory-Elastic Multi-Exit Neural NetworksJiaming Huang, Yi Gao, Wei DongWWW 2024 · 3 citations
- Recurrent Early Exits for Federated Learning with Heterogeneous ClientsRoyson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li et al.ICML 2024 · 13 citations
