STL-SGD: Speeding Up Local SGD with Stagewise Communication Period
Shuheng Shen, Yifei Cheng, Jingchang Liu, Linli Xu
Abstract
Distributed parallel stochastic gradient descent algorithms are workhorses for large scale machine learning tasks. Among them, local stochastic gradient descent (Local SGD) has attracted significant attention due to its low communication complexity. Previous studies prove that the communication complexity of Local SGD with a fixed or an adaptive communication period is in the order of O (N3/2 T1/2) and O (N3/4 T3/4) when the data distributions on clients are identical (IID) or otherwise (Non-IID), where N is the number of clients and T is the number of iterations. In this paper, to accelerate the convergence by reducing the communication complexity, we propose STagewise Local SGD (STL-SGD), which increases the communication period gradually along with decreasing learning rate. We prove that STL-SGD can keep the same convergence rate and linear speedup as mini-batch SGD. In addition, as the benefit of increasing the communication period, when the objective is strongly convex or satisfies the Polyak-Lojasiewicz condition, the communication complexity of STL-SGD is O (N log T ) and O (N1/2 T1/2) for the IID case and the Non-IID case respectively, achieving significant improvements over Local SGD. Experiments on both convex and non-convex problems demonstrate the superior performance of STL-SGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Quadratic Synchronization Rule for Distributed Deep LearningXinran Gu, Kaifeng Lyu, Sanjeev Arora, Jingzhao Zhang et al.ICLR 2024 · 4 citations
- AdaGK-SGD: Adaptive Global Knowledge Guided Distributed Stochastic Gradient DescentHangyu Ye, Weiying Xie, Yunsong Li, Leyuan FangAAAI 2025 · 1 citation
- EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language ModelsJialiang Cheng, Ning Gao, Yun Yue, Zhiling Ye et al.ICLR 2025 · 1 citation
Builds on3
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
Related papers
- Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical NetworksTimothy Castiglia, Anirban Das, Stacy PattersonICLR 2021 · 11 citations
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Federated Learning under Arbitrary Communication PatternsDmitrii Avdiukhin, Shiva Prasad KasiviswanathanICML 2021 · 67 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Convergence of Distributed Adaptive Optimization with Local UpdatesZiheng Cheng, Margalit GlasgowICLR 2025
