Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?
Xuanyu Chen, Nan Yang, Shuai Wang, Dong Yuan
摘要
The recent success of large language models (LLMs) has sparked a growing interest in training large-scale models. As the model size continues to scale, concerns are growing about the depletion of high-quality, well-curated training data. This has led practitioners to explore training approaches like Federated Learning (FL), which can leverage the abundant data on edge devices while maintaining privacy. However, the decentralization of training datasets in FL introduces challenges to scaling large models, a topic that remains under-explored. This paper fills this gap and provides qualitative insights on generalizing the previous model scaling experience to federated learning scenarios. Specifically, we derive a PAC-Bayes (Probably Approximately Correct Bayesian) upper bound for the generalization error of models trained with stochastic algorithms in federated settings and quantify the impact of distributed training data on the optimal model size by finding the analytic solution of model size that minimizes this bound. Our theoretical results demonstrate that the optimal model size has a negative power law relationship with the number of clients if the total training compute is unchanged. Besides, we also find that switching to FL with the same training compute will inevitably reduce the upper bound of generalization performance that the model can achieve through training, and that estimating the optimal model size in federated scenarios should depend on the average training compute across clients. Furthermore, we also empirically validate the correctness of our results with extensive training runs on different models, network settings, and datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- Personalized Federated Learning With Gaussian ProcessesIdan Achituve, Aviv Shamsian, Aviv Navon, Gal Chechik 等NeurIPS 2021 · 被引用 137 次
相关 Paper
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi 等ICML 2021 · 被引用 100 次
- Lessons from Generalization Error Analysis of Federated Learning: You May Communicate Less Often!Milad Sefidgaran, Romain Chor, Abdellatif Zaidi, Yijun WanICML 2024 · 被引用 11 次
- Efficient Device Scheduling with Multi-Job Federated LearningChendi Zhou, Ji Liu, Juncheng Jia, Jingbo Zhou 等AAAI 2022 · 被引用 54 次
- AutoFL: Enabling Heterogeneity-Aware Energy Efficient Federated LearningYoung Geun Kim, Carole-Jean WuMICRO 2021 · 被引用 84 次
- Titanic: Towards Production Federated Learning with Large Language ModelsNingxin Su, Chenghao Hu, Baochun Li, Bo LiINFOCOM 2024 · 被引用 30 次
