Leveraging partial stragglers within gradient coding
Aditya Ramamoorthy, Ruoyu Meng, Vrinda S. Girimaji
摘要
Within distributed learning, workers typically compute gradients on their assigned dataset chunks and send them to the parameter server (PS), which aggregates them to compute either an exact or approximate version of (gradient of the loss function ). However, in large-scale clusters, many workers are slower than their promised speed or even failure-prone. A gradient coding solution introduces redundancy within the assignment of chunks to the workers and uses coding theoretic ideas to allow the PS to recover (exactly or approximately), even in the presence of stragglers. Unfortunately, most existing gradient coding protocols are inefficient from a computation perspective as they coarsely classify workers as operational or failed; the potentially valuable work performed by slow workers (partial stragglers) is ignored. In this work, we present novel gradient coding protocols that judiciously leverage the work performed by partial stragglers. Our protocols are efficient from a computation and communication perspective and numerically stable. For an important class of chunk assignments, we present efficient algorithms for optimizing the relative ordering of chunks within the workers; this ordering affects the overall execution time. For exact gradient reconstruction, our protocol is around faster than the original class of protocols and for approximate gradient reconstruction, the mean-squared-error of our reconstructed gradient is several orders of magnitude better.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Approximate Gradient Coding for Distributed Learning with Heterogeneous StragglersHeekang Song, Wan ChoiNeurIPS 2025 · 被引用 1 次
- Sequential Gradient Coding For Straggler MitigationMuralee Nikhil Krishnan, MohammadReza Ebrahimi, Ashish J. KhistiICLR 2023
- Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient DescentPedro Soto, Ilia Ilmer, Haibin Guan, Jun LiICML 2022 · 被引用 3 次
- Stream Iterative Distributed Coded Computing for Learning Applications in Heterogeneous SystemsHoma Esfahanizadeh, Alejandro Cohen, Muriel MédardINFOCOM 2022 · 被引用 9 次
- Incentive Mechanism Design for Distributed Coded Machine LearningNingning Ding, Zhixuan Fang, Lingjie Duan, Jianwei HuangINFOCOM 2021 · 被引用 20 次
