Is Local SGD Better than Minibatch SGD?
Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai, Brian Bullins, H. Brendan McMahan, Ohad Shamir, Nathan Srebro
Abstract
We study local SGD (also known as parallel SGD and federated averaging), a natural and frequently used stochastic distributed optimization method. Its theoretical foundations are currently lacking and we highlight how all existing error guarantees in the convex setting are dominated by a simple baseline, minibatch SGD. (1) For quadratic objectives we prove that local SGD strictly dominates minibatch SGD and that accelerated local SGD is minimax optimal for quadratics; (2) For general convex objectives we provide the first guarantee that at least sometimes improves over minibatch SGD; (3) We show that indeed local SGD does not dominate minibatch SGD by presenting a lower bound on the performance of local SGD that is worse than the minibatch SGD guarantee.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83656317-ffe3-41db-a5a6-4bfde6f05799Cited by top-tier papers87
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- Clustered Sampling: Low-Variance and Improved Representativity for Clients Selection in Federated LearningYann Fraboni, Richard Vidal, Laetitia Kameni, Marco LorenziICML 2021 · 249 citations
- Preservation of the Global Knowledge by Not-True Distillation in Federated LearningGihun Lee, Minchan Jeong, Yongjin Shin, Sangmin Bae et al.NeurIPS 2022 · 235 citations
- FedSplit: an algorithmic framework for fast federated optimizationReese Pathak, Martin J. WainwrightNeurIPS 2020 · 217 citations
- Federated Accelerated Stochastic Gradient DescentHonglin Yuan, Tengyu MaNeurIPS 2020 · 217 citations
Builds on1
Related papers
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and BeyondChulhee Yun, Shashank Rajput, Suvrit SraICLR 2022 · 47 citations
- Revisiting Consensus Error: A Fine-grained Analysis of Local SGD under Second-order Data HeterogeneityKumar Kshitij Patel, Ali Zindari, Sebastian U. Stich, Lingxiao WangNeurIPS 2025 · 1 citation
- Federated Learning under Arbitrary Communication PatternsDmitrii Avdiukhin, Shiva Prasad KasiviswanathanICML 2021 · 67 citations
- SLowcalSGD : Slow Query Points Improve Local-SGD for Stochastic Convex OptimizationTehila Dahan, Kfir Y. LevyNeurIPS 2024 · 5 citations
