Sequential Gradient Coding For Straggler Mitigation
Muralee Nikhil Krishnan, MohammadReza Ebrahimi, Ashish J. Khisti
摘要
In distributed computing, slower nodes (stragglers) usually become a bottleneck. Gradient Coding (GC), introduced by Tandon et al., is an efficient technique that uses principles of error-correcting codes to distribute gradient computation in the presence of stragglers. In this paper, we consider the distributed computation of a sequence of gradients , where processing of each gradient starts in round- and finishes by round-. Here denotes a delay parameter. For the GC scheme, coding is only across computing nodes and this results in a solution where . On the other hand, having T>0 allows for designing schemes which exploit the temporal dimension as well. In this work, we propose two schemes that demonstrate improved performance compared to GC. Our first scheme combines GC with selective repetition of previously unfinished tasks and achieves improved straggler mitigation. In our second scheme, which constitutes our main contribution, we apply GC to a subset of the tasks and repetition for the remainder of the tasks. We then multiplex these two classes of tasks across workers and rounds in an adaptive manner, based on past straggler patterns. Using theoretical analysis, we demonstrate that our second scheme achieves significant reduction in the computational load. In our experiments, we study a practical setting of concurrently training multiple neural networks over an AWS Lambda cluster involving 256 worker nodes, where our framework naturally applies. We demonstrate that the latter scheme can yield a 16% improvement in runtime over the baseline GC scheme, in the presence of naturally occurring, non-simulated stragglers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Leveraging partial stragglers within gradient codingAditya Ramamoorthy, Ruoyu Meng, Vrinda S. GirimajiNeurIPS 2024 · 被引用 7 次
- Approximate Gradient Coding for Distributed Learning with Heterogeneous StragglersHeekang Song, Wan ChoiNeurIPS 2025 · 被引用 1 次
- Coded Edge ComputingKwang Taik Kim, Carlee Joe-Wong, Mung ChiangINFOCOM 2020 · 被引用 28 次
- Live Gradient Compensation for Evading Stragglers in Distributed LearningJian Xu, Shao-Lun Huang, Linqi Song, Tian LanINFOCOM 2021 · 被引用 29 次
- Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient DescentPedro Soto, Ilia Ilmer, Haibin Guan, Jun LiICML 2022 · 被引用 3 次
