Live Gradient Compensation for Evading Stragglers in Distributed Learning
Jian Xu, Shao-Lun Huang, Linqi Song, Tian Lan
Abstract
The training efficiency of distributed learning systems is vulnerable to stragglers, namely, those slow worker nodes. A naive strategy is performing the distributed learning by incor-porating the fastest K workers and ignoring these stragglers, which may induce high deviation for non-IID data. To tackle this, we develop a Live Gradient Compensation (LGC) strategy to incorporate the one-step delayed gradients from stragglers, aiming to accelerate learning process and utilize the stragglers simultaneously. In LGC framework, mini-batch data are divided into smaller blocks and processed separately, which makes the gradient computed based on partial work accessible. In addition, we provide theoretical convergence analysis of our algorithm for non-convex optimization problem under non-IID training data to show that LGC-SGD has almost the same convergence error as full synchronous SGD. The theoretical results also allow us to quantify a novel tradeoff in minimizing training time and error by selecting the optimal straggler threshold. Finally, extensive simulation experiments of image classification on CIFAR-10 dataset are conducted, and the numerical results demonstrate the effectiveness of our proposed strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44ef5005-e7a3-4515-94dd-7096ac81e83aCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Approximate Gradient Coding for Distributed Learning with Heterogeneous StragglersHeekang Song, Wan ChoiNeurIPS 2025 · 1 citation
- Leveraging partial stragglers within gradient codingAditya Ramamoorthy, Ruoyu Meng, Vrinda S. GirimajiNeurIPS 2024 · 7 citations
- DAGC: Data-Aware Adaptive Gradient CompressionRongwei Lu, Jiajun Song, Bin Chen, Laizhong Cui et al.INFOCOM 2023 · 12 citations
- Sequential Gradient Coding For Straggler MitigationMuralee Nikhil Krishnan, MohammadReza Ebrahimi, Ashish J. KhistiICLR 2023
- Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical NetworksTimothy Castiglia, Anirban Das, Stacy PattersonICLR 2021 · 11 citations
