Beyond spectral gap: the role of the topology in decentralized learning
Thijs Vogels, Hadrien Hendrikx, Martin Jaggi
摘要
In data-parallel optimization of machine learning models, workers collaborate to improve their estimates of the model: more accurate gradients allow them to use larger learning rates and optimize faster. We consider the setting in which all workers sample from the same dataset, and communicate over a sparse graph (decentralized). In this setting, current theory fails to capture important aspects of real-world behavior. First, the 'spectral gap' of the communication graph is not predictive of its empirical performance in (deep) learning. Second, current theory does not explain that collaboration enables larger learning rates than training alone. In fact, it prescribes smaller learning rates, which further decrease as graphs become larger, failing to explain convergence in infinite graphs. This paper aims to paint an accurate picture of sparsely-connected distributed optimization when workers share the same data distribution. We quantify how the graph topology influences convergence in a quadratic toy problem and provide theoretical results for general smooth and (strongly) convex objectives. Our theory matches empirical observations in deep learning, and accurately describes the relative merits of different graph topologies. Code: github.com/epfml/topology-in-decentralized-learning
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Topology-aware Generalization of Decentralized SGDTongtian Zhu, Fengxiang He, Lan Zhang, Zhengyang Niu 等ICML 2022 · 被引用 58 次
- DePRL: Achieving Linear Convergence Speedup in Personalized Decentralized Learning with Shared RepresentationsGuojun Xiong, Gang Yan, Shiqiang Wang, Jian LiAAAI 2024 · 被引用 9 次
- On The Surprising Effectiveness of a Single Global Merging in Decentralized LearningTongtian Zhu, Tianyu Zhang, Mingze Wang, Zhanpeng Zhou 等ICLR 2026 · 被引用 2 次
- Structured Cooperative Learning with Graphical Model PriorsShuangtong Li, Tianyi Zhou, Xinmei Tian, Dacheng TaoICML 2023
- DICE: Data Influence Cascade in Decentralized LearningTongtian Zhu, Wenhao Li, Can Wang, Fengxiang HeICLR 2025
它引用的顶会 Paper5
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Exponential Graph is Provably Efficient for Decentralized Deep TrainingBicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu 等NeurIPS 2021 · 被引用 123 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- RelaySum for Decentralized Deep Learning on Heterogeneous DataThijs Vogels, Lie He, Anastasia Koloskova, Sai Praneeth Karimireddy 等NeurIPS 2021 · 被引用 78 次
相关 Paper
- Improved Convergence Analysis of Topology Dependence in Decentralized SGDYuki Takezawa, Anastasiia Koloskova, Sebastian StichICML 2026
- Scalable Decentralized Learning with TeleportationYuki Takezawa, Sebastian U. StichICLR 2025
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi 等ICML 2021 · 被引用 100 次
- Communication-Efficient Topologies for Decentralized Learning with Consensus RateZhuoqing Song, Weijian Li, Kexin Jin, Lei Shi 等NeurIPS 2022 · 被引用 48 次
- An Improved Analysis of Gradient Tracking for Decentralized Machine LearningAnastasia Koloskova, Tao Lin, Sebastian U. StichNeurIPS 2021 · 被引用 148 次
