On the Importance and Applicability of Pre-Training for Federated Learning
Hong-You Chen, Cheng-Hao Tu, Ziwei Li, Han-Wei Shen, Wei-Lun Chao
Abstract
Pre-training is prevalent in nowadays deep learning to improve the learned model's performance. However, in the literature on federated learning (FL), neural networks are mostly initialized with random weights. These attract our interest in conducting a systematic study to explore pre-training for FL. Across multiple visual recognition benchmarks, we found that pre-training can not only improve FL, but also close its accuracy gap to the counterpart centralized learning, especially in the challenging cases of non-IID clients' data. To make our findings applicable to situations where pre-trained models are not directly available, we explore pre-training with synthetic data or even with clients' data in a decentralized manner, and found that they can already improve FL notably. Interestingly, many of the techniques we explore are complementary to each other to further boost the performance, and we view this as a critical result toward scaling up deep FL for real-world applications. We conclude our paper with an attempt to understand the effect of pre-training on FL. We found that pre-training enables the learned global models under different clients' data conditions to converge to the same loss basin, and makes global aggregation in FL more stable. Nevertheless, pre-training seems to not alleviate local model drifting, a fundamental problem in FL under non-IID data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66f59be0-9ba1-40f8-8ef6-edd385461a23Cited by top-tier papers17
- Efficient Model Personalization in Federated Learning via Client-Specific Prompt GenerationFu-En Yang, Chien-Yi Wang, Yu-Chiang Frank WangICCV 2023 · 112 citations
- Enhancing One-Shot Federated Learning Through Data and Ensemble Co-BoostingRong Dai, Yonggang Zhang, Ang Li, Tongliang Liu et al.ICLR 2024 · 40 citations
- Guiding The Last Layer in Federated Learning with Pre-Trained ModelsGwen Legate, Nicolas Bernier, Lucas Page-Caccia, Edouard Oyallon et al.NeurIPS 2023 · 31 citations
- Layer-wise linear mode connectivityLinara Adilova, Maksym Andriushchenko, Michael Kamp, Asja Fischer et al.ICLR 2024 · 22 citations
- Recurrent Early Exits for Federated Learning with Heterogeneous ClientsRoyson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li et al.ICML 2024 · 13 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- Exploiting Label Skews in Federated Learning with Model ConcatenationYiqun Diao, Qinbin Li, Bingsheng HeAAAI 2024 · 39 citations
- Rethinking the Starting Point: Collaborative Pre-Training for Federated Downstream TasksYun-Wei Chu, Dong-Jun Han, Seyyedali Hosseinalipour, Christopher G. BrintonAAAI 2025 · 1 citation
- DYNAFED: Tackling Client Data Heterogeneity with Global DynamicsRenjie Pi, Weizhong Zhang, Yueqi Xie, Jiahui Gao et al.CVPR 2023
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp et al.ICLR 2021 · 1,166 citations
- FedDC: Federated Learning with Non-IID Data via Local Drift Decoupling and CorrectionLiang Gao, Huazhu Fu, Li Li, Yingwen Chen et al.CVPR 2022 · 307 citations
