An Improved Analysis of Gradient Tracking for Decentralized Machine Learning
Anastasia Koloskova, Tao Lin, Sebastian U. Stich
摘要
We consider decentralized machine learning over a network where the training data is distributed across agents, each of which can compute stochastic model updates on their local data. The agent's common goal is to find a model that minimizes the average of all local loss functions. While gradient tracking (GT) algorithms can overcome a key challenge, namely accounting for differences between workers' local data distributions, the known convergence rates for GT algorithms are not optimal with respect to their dependence on the mixing parameter (related to the spectral gap of the connectivity matrix). We provide a tighter analysis of the GT method in the stochastic strongly convex, convex and non-convex settings. We improve the dependency on from to in the noiseless case and from to in the general stochastic case, where is related to the negative eigenvalues of the connectivity matrix (and is a constant in most practical applications). This improvement was possible due to a new proof technique which could be of independent interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 被引用 200 次
- BEER: Fast Rate for Decentralized Nonconvex Optimization with Communication CompressionHaoyu Zhao, Boyue Li, Zhize Li, Peter Richtárik 等NeurIPS 2022 · 被引用 76 次
- Lower Bounds and Nearly Optimal Algorithms in Distributed Learning with Communication CompressionXinmeng Huang, Yiming Chen, Wotao Yin, Kun YuanNeurIPS 2022 · 被引用 49 次
- Communication-Efficient Topologies for Decentralized Learning with Consensus RateZhuoqing Song, Weijian Li, Kexin Jin, Lei Shi 等NeurIPS 2022 · 被引用 48 次
- Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized OptimizationKun Yuan, Xinmeng Huang, Yiming Chen, Xiaohan Zhang 等NeurIPS 2022 · 被引用 40 次
它引用的顶会 Paper7
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Quasi-global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous DataTao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, Martin JaggiICML 2021 · 被引用 118 次
- Consensus Control for Decentralized Deep LearningLingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi 等ICML 2021 · 被引用 100 次
- RelaySum for Decentralized Deep Learning on Heterogeneous DataThijs Vogels, Lie He, Anastasia Koloskova, Sai Praneeth Karimireddy 等NeurIPS 2021 · 被引用 78 次
相关 Paper
- Improving the Sample and Communication Complexity for Decentralized Non-Convex Optimization: Joint Gradient Estimation and TrackingHaoran Sun, Songtao Lu, Mingyi HongICML 2020 · 被引用 57 次
- Compressed Decentralized Proximal Stochastic Gradient Method for Nonconvex Composite Problems with Heterogeneous DataYonggui Yan, Jie Chen, Pin-Yu Chen, Xiaodong Cui 等ICML 2023 · 被引用 18 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- Decentralized Riemannian Gradient Descent on the Stiefel ManifoldShixiang Chen, Alfredo García, Mingyi Hong, Shahin ShahrampourICML 2021 · 被引用 64 次
- Taming Subnet-Drift in D2D-Enabled Fog Learning: A Hierarchical Gradient Tracking ApproachEvan Chen, Shiqiang Wang, Christopher G. BrintonINFOCOM 2024 · 被引用 5 次
