Accelerating Gossip SGD with Periodic Global Averaging
Yiming Chen, Kun Yuan, Yingya Zhang, Pan Pan, Yinghui Xu, Wotao Yin
Abstract
Communication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient than the prevalent parallel SGD. However, its convergence rate is reversely proportional to quantity which measures the network connectivity. On large and sparse networks where , Gossip SGD requires more iterations to converge, which offsets against its communication benefit. This paper introduces Gossip-PGA, which adds Periodic Global Averaging into Gossip SGD. Its transient stage, i.e., the iterations required to reach asymptotic linear speedup stage, improves from to for non-convex problems. The influence of network topology in Gossip-PGA can be controlled by the averaging period . Its transient-stage complexity is also superior to Local SGD which has order . Empirical results of large-scale training on image classification (ResNet50) and language modeling (BERT) validate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu et al.NeurIPS 2021 · 123 citations
- Lower Bounds and Nearly Optimal Algorithms in Distributed Learning with Communication CompressionXinmeng Huang, Yiming Chen, Wotao Yin, Kun YuanNeurIPS 2022 · 49 citations
- Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized OptimizationKun Yuan, Xinmeng Huang, Yiming Chen, Xiaohan Zhang et al.NeurIPS 2022 · 40 citations
- Epidemic Learning: Boosting Decentralized Learning with Randomized CommunicationMartijn de Vos, Sadegh Farhadkhani, Rachid Guerraoui, Anne-Marie Kermarrec et al.NeurIPS 2023 · 39 citations
- Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time ConvergenceYuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa et al.NeurIPS 2023 · 22 citations
Builds on6
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 263 citations
Related papers
- Scalable Decentralized Learning with TeleportationYuki Takezawa, Sebastian U. StichICLR 2025
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 57 citations
- DSGD-CECA: Decentralized SGD with Communication-Optimal Exact Consensus AlgorithmLisang Ding, Kexin Jin, Bicheng Ying, Kun Yuan et al.ICML 2023 · 12 citations
- Learn Locally, Correct Globally: A Distributed Algorithm for Training Graph Neural NetworksMorteza Ramezani, Weilin Cong, Mehrdad Mahdavi, Mahmut T. Kandemir et al.ICLR 2022 · 35 citations
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu et al.ICML 2023 · 7 citations
