Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced Topology
Yan Huang, Ying Sun, Zehan Zhu, Changzhi Yan, Jinming Xu
摘要
We develop a general framework unifying several gradient-based stochastic optimization methods for empirical risk minimization problems both in centralized and distributed scenarios. The framework hinges on the introduction of an augmented graph consisting of nodes modeling the samples and edges modeling both the inter-device communication and intra-device stochastic gradient computation. By designing properly the topology of the augmented graph, we are able to recover as special cases the renowned Local-SGD and DSGD algorithms, and provide a unified perspective for variance-reduction (VR) and gradient-tracking (GT) methods such as SAGA, Local-SVRG and GT-SAGA. We also provide a unified convergence analysis for smooth and (strongly) convex objectives relying on a proper structured Lyapunov function, and the obtained rate can recover the best known results for many existing algorithms. The rate results further reveal that VR and GT methods can effectively eliminate data heterogeneity within and across devices, respectively, enabling the exact convergence of the algorithm to the optimal solution. Numerical experiments confirm the findings in this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Taming Subnet-Drift in D2D-Enabled Fog Learning: A Hierarchical Gradient Tracking ApproachEvan Chen, Shiqiang Wang, Christopher G. BrintonINFOCOM 2024 · 被引用 5 次
- Reducing Training Time in Cross-Silo Federated Learning using Multigraph TopologyTuong Do, Binh X. Nguyen, Vuong Pham, Toan Tran 等ICCV 2023 · 被引用 4 次
- Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD MethodsAlexander Tyurin, Danil SivtsovICLR 2026 · 被引用 2 次
- Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive StepsizesYan Huang, Xiang Li, Yipeng Shen, Niao He 等NeurIPS 2024 · 被引用 2 次
- Structured Cooperative Learning with Graphical Model PriorsShuangtong Li, Tianyi Zhou, Xinmei Tian, Dacheng TaoICML 2023
它引用的顶会 Paper6
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- The Non-IID Data Quagmire of Decentralized Machine LearningKevin Hsieh, Amar Phanishayee, Onur Mutlu, Phillip B. GibbonsICML 2020 · 被引用 672 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Improving the Sample and Communication Complexity for Decentralized Non-Convex Optimization: Joint Gradient Estimation and TrackingHaoran Sun, Songtao Lu, Mingyi HongICML 2020 · 被引用 57 次
- Accelerating Gossip SGD with Periodic Global AveragingYiming Chen, Kun Yuan, Yingya Zhang, Pan Pan 等ICML 2021 · 被引用 49 次
相关 Paper
- A Short and Unified Convergence Analysis of the SAG, SAGA, and IAG AlgorithmsFeng Zhu, Robert Heath, Aritra MitraICML 2026
- Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionKenta Niwa, Guoqiang Zhang, W. Bastiaan Kleijn, Noboru Harada 等ICML 2021 · 被引用 16 次
- A framework for bilevel optimization that enables stochastic and global variance reduction algorithmsMathieu Dagréou, Pierre Ablin, Samuel Vaiter, Thomas MoreauNeurIPS 2022 · 被引用 149 次
- Efficient Sign-Based Optimization: Accelerating Convergence via Variance ReductionWei Jiang, Sifan Yang, Wenhao Yang, Lijun ZhangNeurIPS 2024 · 被引用 19 次
- Escaping Saddle Points with Bias-Variance Reduced Local Perturbed SGD for Communication Efficient Nonconvex Distributed LearningTomoya Murata, Taiji SuzukiNeurIPS 2022 · 被引用 4 次
