Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks
Timothy Castiglia, Anirban Das, Stacy Patterson
Abstract
We propose Multi-Level Local SGD, a distributed stochastic gradient method for learning a smooth, non-convex objective in a multi-level communication network with heterogeneous workers. Our network model consists of a set of disjoint sub-networks, with a single hub and multiple workers; further, workers may have different operating rates. The hubs exchange information with one another via a connected, but not necessarily complete communication network. In our algorithm, sub-networks execute a distributed SGD algorithm, using a hub-and-spoke paradigm, and the hubs periodically average their models with neighboring hubs. We first provide a unified mathematical framework that describes the Multi-Level Local SGD algorithm. We then present a theoretical analysis of the algorithm; our analysis shows the dependence of the convergence error on the worker node heterogeneity, hub network topology, and the number of local, sub-network, and global iterations. We illustrate the effectiveness of our algorithm in a multi-level network with slow workers via simulation-based experiments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1b1f84bf-da5b-4805-b419-1651ffbd0540Cited by top-tier papers5
- Hierarchical Federated Learning with Multi-Timescale Gradient CorrectionWenzhi Fang, Dong-Jun Han, Evan Chen, Shiqiang Wang et al.NeurIPS 2024 · 34 citations
- Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced TopologyYan Huang, Ying Sun, Zehan Zhu, Changzhi Yan et al.ICML 2022 · 18 citations
- EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language ModelsJialiang Cheng, Ning Gao, Yun Yue, Zhiling Ye et al.ICLR 2025 · 1 citation
- Unified Analyses for Hierarchical Federated Learning: Topology Selection under Data HeterogeneityZiyi Zhou, Yipeng Li, Xinchen LyuICLR 2026
- HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model TrainingGeon-Woo Kim, Junbo Li, Shashidhar Gandham, Omar Baldonado et al.ICML 2025
Related papers
- Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGDJiayi Wang, Shiqiang Wang, Rong-Rong Chen, Mingyue JiAAAI 2022 · 81 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Federated Learning under Arbitrary Communication PatternsDmitrii Avdiukhin, Shiva Prasad KasiviswanathanICML 2021 · 67 citations
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated LearningTomoya Murata, Taiji SuzukiICML 2021 · 61 citations
