DarkDistill: Difficulty-Aligned Federated Early-Exit Network Training on Heterogeneous Devices
Lehao Qu, Shuyuan Li, Zimu Zhou, Boyi Liu, Yi Xu, Yongxin Tong
摘要
Early-exit networks (EENs), which adapt their computational depths based on input samples, are widely adopted to accelerate inference in edge computing applications. The effectiveness of EENs relies on difficulty-aware training, which tailors shallow exits for simple samples and deep exits for complex ones. However, existing difficulty-aware training schemes assume centralized environments with sufficient data, which become invalid with real-world edge devices. In this paper, we explore difficulty-aware training in a federated manner, where EENs are collaboratively trained on heterogeneous devices. We observe the cross-model exit unalignment phenomenon, a unique problem when aggregating local EENs into a cohesive global model. To address this problem, we design a novel Difficulty-Aligned Reverse Knowledge Distillation scheme named DarkDistill that preserves the difficulty-specific specialization for aggregating heterogeneous local models. Instead of direct parameter averaging, it trains difficulty-conditional data generators, and selectively transfers generated knowledge of specific difficulty among matched exits of heterogeneous EENs. Evaluations show that Dark-Distill outperforms the state-of-the-arts in both full-parameter and parameter-efficient fine-tuning of EENs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language ModelsYunlong Chu, Minglai Shao, Yuhang Liu, Bing Hao 等KDD 2026 · 被引用 2 次
- Towards Asynchronous Client Collaboration in Personalized Federated LearningBoyi Liu, Zimu Zhou, Pengfei Gao, Shuo Kang 等INFOCOM 2026
它引用的顶会 Paper25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 被引用 957 次
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningLin Zhang, Li Shen, Liang Ding, Dacheng Tao 等CVPR 2022 · 被引用 339 次
相关 Paper
- ScaleFL: Resource-Adaptive Federated Learning with Heterogeneous ClientsFatih Ilhan, Gong Su, Ling LiuCVPR 2023
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 被引用 205 次
- Harmonized Dense Knowledge Distillation Training for Multi-Exit ArchitecturesXinglu Wang, Yingming LiAAAI 2021 · 被引用 26 次
- Unlocking the Non-deterministic Computing Power with Memory-Elastic Multi-Exit Neural NetworksJiaming Huang, Yi Gao, Wei DongWWW 2024 · 被引用 3 次
- Recurrent Early Exits for Federated Learning with Heterogeneous ClientsRoyson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li 等ICML 2024 · 被引用 13 次
