Where to Begin? On the Impact of Pre-Training and Initialization in Federated Learning
John Nguyen, Jianyu Wang, Kshitiz Malik, Maziar Sanjabi, Michael G. Rabbat
Abstract
An oft-cited challenge of federated learning is the presence of heterogeneity. Data heterogeneity refers to the fact that data from different clients may follow very different distributions. System heterogeneity refers to the fact that client devices have different system capabilities. A considerable number of federated optimization methods address this challenge. In the literature, empirical evaluations usually start federated training from random initialization. However, in many practical applications of federated learning, the server has access to proxy data for the training task that can be used to pre-train a model before starting federated training. We empirically study the impact of starting from a pre-trained model in federated learning using four standard federated learning benchmark datasets. Unsurprisingly, starting from a pre-trained model reduces the training time required to reach a target error rate and enables the training of more accurate models (up to 40%) than is possible when starting from random initialization. Surprisingly, we also find that starting federated learning from a pre-trained initialization reduces the effect of both data and system heterogeneity. We recommend that future work proposing and evaluating federated optimization methods evaluate the performance when starting from random and pre-trained initializations. We also believe this study raises several questions for further work on understanding the role of heterogeneity in federated optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0ef861c-0d88-40d1-b3a5-6f1e19df69fcCited by top-tier papers15
- Efficient Model Personalization in Federated Learning via Client-Specific Prompt GenerationFu-En Yang, Chien-Yi Wang, Yu-Chiang Frank WangICCV 2023 · 112 citations
- FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language ModelsJingwei Sun, Ziyue Xu, Hongxu Yin, Dong Yang et al.ICML 2024 · 38 citations
- Guiding The Last Layer in Federated Learning with Pre-Trained ModelsGwen Legate, Nicolas Bernier, Lucas Page-Caccia, Edouard Oyallon et al.NeurIPS 2023 · 31 citations
- Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced DataPei-Yau Weng, Minh Hoang, Lam M. Nguyen, My T. Thai et al.NeurIPS 2024 · 19 citations
- Recurrent Early Exits for Federated Learning with Heterogeneous ClientsRoyson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li et al.ICML 2024 · 13 citations
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
Related papers
- Improved Modelling of Federated Datasets using Mixtures-of-Dirichlet-MultinomialsJonathan Scott, Áine CahillICML 2024 · 2 citations
- Client2Vec: Improving Federated Learning by Distribution Shifts Aware Client IndexingYongxin Guo, Lin Wang, Xiaoying Tang, Tao LinICCV 2025
- When Do Curricula Work in Federated Learning?Saeed Vahidian, Sreevatsank Kadaveru, Woonjoon Baek, Weijia Wang et al.ICCV 2023 · 12 citations
- TiFL: A Tier-based Federated Learning SystemZheng Chai, Ahsan Ali, Syed Zawad, Stacey Truex et al.HPDC 2020 · 330 citations
- FedProto: Federated Prototype Learning across Heterogeneous ClientsYue Tan, Guodong Long, Lu Liu, Tianyi Zhou et al.AAAI 2022 · 851 citations
