FedAvg Converges to Zero Training Loss Linearly for Overparameterized Multi-Layer Neural Networks
Bingqing Song, Prashant Khanduri, Xinwei Zhang, Jinfeng Yi, Mingyi Hong
摘要
Federated Learning (FL) is a distributed learning paradigm that allows multiple clients to learn a joint model by utilizing privately held data at each client. Significant research efforts have been devoted to develop advanced algorithms that deal with the situation where the data at individual clients have heterogeneous distributions. In this work, we show that data heterogeneity can be dealt from a different perspective. That is, by utilizing a certain overparameterized multi-layer neural network at each client, even the vanilla FedAvg (a.k.a. the Local SGD) algorithm can accurately optimize the training problem: When each client has a neural network with one wide layer of size N (where N is the number of total training samples), followed by layers of smaller widths, FedAvg converges linearly to a solution that achieves (almost) zero training loss, without requiring any assumptions on the clients' data distributions. To our knowledge, this is the first work that demonstrates such resilience to data heterogeneity for FedAvg when trained on multi-layer neural networks. Our experiments also confirm that, neural networks of large size can achieve better and more stable performance for FL problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Convergence and Generalization in Federated Learning through Feature Learning TheoryWei Huang, Ye Shi, Zhongyi Cai, Taiji SuzukiICLR 2024 · 被引用 17 次
- Enhancing Network Attack Detection with Distributed and In-Network Data Collection SystemSeyed Mohammad Mehdi Mirnajafizadeh, Ashwin Raam Sethuram, David Mohaisen, DaeHun Nyang 等USENIX Security 2024 · 被引用 12 次
- Widening the Network Mitigates the Impact of Data Heterogeneity on FedAvgLike Jian, Dong LiuICML 2025
它引用的顶会 Paper12
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated LearningHaibo Yang, Minghong Fang, Jia LiuICLR 2021 · 被引用 310 次
相关 Paper
- On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous DataBo Li, Mikkel N. Schmidt, Tommy S. Alstrøm, Sebastian U. StichCVPR 2023
- Understanding Clipping for Federated Learning: Convergence and Client-Level Differential PrivacyXinwei Zhang, Xiangyi Chen, Mingyi Hong, Steven Wu 等ICML 2022 · 被引用 134 次
- Elastic Aggregation for Federated OptimizationDengsheng Chen, Jie Hu, Vince Junkai Tan, Xiaoming Wei 等CVPR 2023
- Architecture Agnostic Federated Learning for Neural NetworksDisha Makhija, Xing Han, Nhat Ho, Joydeep GhoshICML 2022 · 被引用 62 次
- Local Learning Matters: Rethinking Data Heterogeneity in Federated LearningMatías Mendieta, Taojiannan Yang, Pu Wang, Minwoo Lee 等CVPR 2022 · 被引用 176 次
