Towards Model Agnostic Federated Learning Using Knowledge Distillation
Andrei Afonin, Sai Praneeth Karimireddy
Abstract
Is it possible to design an universal API for federated learning using which an ad-hoc group of data-holders (agents) collaborate with each other and perform federated learning? Such an API would necessarily need to be model-agnostic i.e. make no assumption about the model architecture being used by the agents, and also cannot rely on having representative public data at hand. Knowledge distillation (KD) is the obvious tool of choice to design such protocols. However, surprisingly, we show that most natural KD-based federated learning protocols have poor performance. To investigate this, we propose a new theoretical framework, Federated Kernel ridge regression, which can capture both model heterogeneity as well as data heterogeneity. Our analysis shows that the degradation is largely due to a fundamental limitation of knowledge distillation under data heterogeneity. We further validate our framework by analyzing and designing new protocols based on KD. Their performance on real world experiments using neural networks, though still unsatisfactory, closely matches our theoretical predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0af340f2-96f1-40f5-9dd8-495f8547bca2Cited by top-tier papers12
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li et al.NeurIPS 2023 · 64 citations
- Resource-Adaptive Federated Learning with All-In-One Neural CompositionYiqun Mei, Pengfei Guo, Mo Zhou, Vishal PatelNeurIPS 2022 · 62 citations
- Accelerated Federated Learning with Decoupled Adaptive OptimizationJiayin Jin, Jiaxiang Ren, Yang Zhou, Lingjuan Lyu et al.ICML 2022 · 62 citations
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent KernelsYaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma et al.NeurIPS 2022 · 38 citations
- Agglomerative Federated Learning: Empowering Larger Model Training via End-Edge-Cloud CollaborationZhiyuan Wu, Sheng Sun, Yuwei Wang, Min Liu et al.INFOCOM 2024 · 29 citations
Builds on7
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
Related papers
- Towards Understanding Ensemble Distillation in Federated LearningSejun Park, Kihun Hong, Ganguk HwangICML 2023 · 9 citations
- FedGMKD: An Efficient Prototype Federated Learning Framework through Knowledge Distillation and Discrepancy-Aware AggregationJianqiao Zhang, Caifeng Shan, Jungong HanNeurIPS 2024 · 35 citations
- FedCD: Towards Consolidated Distillation for Heterogeneous Federated LearningYichen Li, Hang Su, Huifa Li, Haolin Yang et al.AAAI 2026
- FedAKD: Federated Adaptive Knowledge Distillation via Global Knowledge Calibration and DecouplingYingchao Wang, Wenqi Niu, Hanpo HouWWW 2026
- A Hierarchical Knowledge Transfer Framework for Heterogeneous Federated LearningYongheng Deng, Ju Ren, Cheng Tang, Feng Lyu et al.INFOCOM 2023 · 37 citations
