B-ary Tree Push-Pull Method is Provably Efficient for Distributed Learning on Heterogeneous Data
Runze You, Shi Pu
摘要
This paper considers the distributed learning problem where a group of agents cooperatively minimizes the summation of their local cost functions based on peer-to-peer communication. Particularly, we propose a highly efficient algorithm, termed "B-ary Tree Push-Pull" (BTPP), that employs two B-ary spanning trees for distributing the information related to the parameters and stochastic gradients across the network. The simple method is efficient in communication since each agent interacts with at most (B + 1) neighbors per iteration. More importantly, BTPP achieves linear speedup for smooth nonconvex and strongly convex objective functions with only Õ(n) and Õ(1) transient iterations, respectively, significantly outperforming the state-of-the-art results to the best of our knowledge. Our code is available at https://github.com/ryou98/BTPP . ALGORITHM PER-ITER COMM. SIZE n BASED GRAPH TRANS. ITER. 1) ARBITRARY 2 Õ(n) ALGORITHM PER-ITER COMM. SIZE n BASED GRAPH TRANS. ITER. DSGD (RING) [18] Θ(1) ARBITRARY 1 Õ(n 5 ) STATIC EXP. [34] Θ(ln(n)) ARBITRARY 1 Õ(n) O.-P. EXP. [34] 1 POWER OF 2 Θ(ln(n)) Õ(n) RELAYSGD [31] Θ(1)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- An Improved Analysis of Gradient Tracking for Decentralized Machine LearningAnastasia Koloskova, Tao Lin, Sebastian U. StichNeurIPS 2021 · 被引用 148 次
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu 等NeurIPS 2021 · 被引用 123 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- RelaySum for Decentralized Deep Learning on Heterogeneous DataThijs Vogels, Lie He, Anastasia Koloskova, Sai Praneeth Karimireddy 等NeurIPS 2021 · 被引用 78 次
相关 Paper
- Low Sample and Communication Complexities in Decentralized Learning: A Triple Hybrid ApproachXin Zhang, Jia Liu, Zhengyuan Zhu, Elizabeth Serena BentleyINFOCOM 2021 · 被引用 6 次
- Decentralized Gossip-Based Stochastic Bilevel Optimization over Communication NetworksShuoguang Yang, Xuezhou Zhang, Mengdi WangNeurIPS 2022 · 被引用 66 次
- Epidemic Learning: Boosting Decentralized Learning with Randomized CommunicationMartijn de Vos, Sadegh Farhadkhani, Rachid Guerraoui, Anne-Marie Kermarrec 等NeurIPS 2023 · 被引用 39 次
- STL-SGD: Speeding Up Local SGD with Stagewise Communication PeriodShuheng Shen, Yifei Cheng, Jingchang Liu, Linli XuAAAI 2021 · 被引用 12 次
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
