Coded Computing for Resilient Distributed Computing: A Learning-Theoretic Framework
Parsa Moradi, Behrooz Tahmasebi, Mohammad Ali Maddah-Ali
摘要
Coded computing has emerged as a promising framework for tackling significant challenges in large-scale distributed computing, including the presence of slow, faulty, or compromised servers. In this approach, each worker node processes a combination of the data, rather than the raw data itself. The final result then is decoded from the collective outputs of the worker nodes. However, there is a significant gap between current coded computing approaches and the broader landscape of general distributed computing, particularly when it comes to machine learning workloads. To bridge this gap, we propose a novel foundation for coded computing, integrating the principles of learning theory, and developing a framework that seamlessly adapts with machine learning applications. In this framework, the objective is to find the encoder and decoder functions that minimize the loss function, defined as the mean squared error between the estimated and true values. Facilitating the search for the optimum decoding and functions, we show that the loss function can be upper-bounded by the summation of two terms: the generalization error of the decoding function and the training error of the encoding function. Focusing on the second-order Sobolev space, we then derive the optimal encoder and decoder. We show that in the proposed solution, the mean squared error of the estimation decays with the rate of and in noiseless and noisy computation settings, respectively, where is the number of worker nodes with at most slow servers (stragglers). Finally, we evaluate the proposed scheme on inference tasks for various machine learning models and demonstrate that the proposed framework outperforms the state-of-the-art in terms of accuracy and rate of convergence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A Scalable Approach for Privacy-Preserving Collaborative Machine LearningJinhyun So, Basak Güler, Salman AvestimehrNeurIPS 2020 · 被引用 60 次
- ApproxIFER: A Model-Agnostic Approach to Resilient and Robust Prediction Serving SystemsMahdi Soleymani, Ramy E. Ali, Hessam Mahdavifar, Amir Salman AvestimehrAAAI 2022 · 被引用 10 次
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han 等CVPR 2021
相关 Paper
- Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient DescentPedro Soto, Ilia Ilmer, Haibin Guan, Jun LiICML 2022 · 被引用 3 次
- Leveraging partial stragglers within gradient codingAditya Ramamoorthy, Ruoyu Meng, Vrinda S. GirimajiNeurIPS 2024 · 被引用 7 次
- Coded Edge ComputingKwang Taik Kim, Carlee Joe-Wong, Mung ChiangINFOCOM 2020 · 被引用 28 次
- Incentive Mechanism Design for Distributed Coded Machine LearningNingning Ding, Zhixuan Fang, Lingjie Duan, Jianwei HuangINFOCOM 2021 · 被引用 20 次
- Stream Iterative Distributed Coded Computing for Learning Applications in Heterogeneous SystemsHoma Esfahanizadeh, Alejandro Cohen, Muriel MédardINFOCOM 2022 · 被引用 9 次
