Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
Tehila Dahan, Kfir Y. Levy
摘要
We address the challenges of Byzantine-robust training in asynchronous distributed machine learning systems, aiming to enhance efficiency amid massive parallelization and heterogeneous computing resources. Asynchronous systems, marked by independently operating workers and intermittent updates, uniquely struggle with maintaining integrity against Byzantine failures, which encompass malicious or erroneous actions that disrupt learning. The inherent delays in such settings not only introduce additional bias to the system but also obscure the disruptions caused by Byzantine faults. To tackle these issues, we adapt the Byzantine framework to asynchronous dynamics by introducing a novel weighted robust aggregation framework. This allows for the extension of robust aggregators and a recent meta-aggregator to their weighted versions, mitigating the effects of delayed updates. By further incorporating a recent variance-reduction technique, we achieve an optimal convergence rate for the first time in an asynchronous Byzantine environment. Our methodology is rigorously validated through empirical and theoretical analysis, demonstrating its effectiveness in enhancing fault tolerance and optimizing performance in asynchronous ML systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 被引用 247 次
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 被引用 192 次
- Zeno++: Robust Fully Asynchronous SGDCong Xie, Sanmi Koyejo, Indranil GuptaICML 2020 · 被引用 137 次
- Byzantine Machine Learning Made Easy By Resilient Averaging of MomentumsSadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot 等ICML 2022 · 被引用 96 次
- Asynchronous SGD Beats Minibatch SGD Under Arbitrary DelaysKonstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. WoodworthNeurIPS 2022 · 被引用 95 次
相关 Paper
- Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous TrainingTehila Dahan, Kfir Yehuda LevyICML 2024 · 被引用 3 次
- Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex OptimizationHamidreza Almasi, Harsh Mishra, Balajee Vamanan, Sathya N. RaviICLR 2024
- Variance Reduction is an Antidote to Byzantines: Better Rates, Weaker Assumptions and Communication Compression as a Cherry on the TopEduard Gorbunov, Samuel Horváth, Peter Richtárik, Gauthier GidelICLR 2023 · 被引用 5 次
- Byzantine-Tolerant Methods for Distributed Variational InequalitiesNazarii Tupitsa, Abdulla Jasem Almansoori, Yanlin Wu, Martin Takác 等NeurIPS 2023 · 被引用 3 次
- Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine WorkersRon Dorfman, Naseem Yehya, Kfir Yehuda LevyICML 2024 · 被引用 5 次
