Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
Tehila Dahan, Kfir Y. Levy
Abstract
We address the challenges of Byzantine-robust training in asynchronous distributed machine learning systems, aiming to enhance efficiency amid massive parallelization and heterogeneous computing resources. Asynchronous systems, marked by independently operating workers and intermittent updates, uniquely struggle with maintaining integrity against Byzantine failures, which encompass malicious or erroneous actions that disrupt learning. The inherent delays in such settings not only introduce additional bias to the system but also obscure the disruptions caused by Byzantine faults. To tackle these issues, we adapt the Byzantine framework to asynchronous dynamics by introducing a novel weighted robust aggregation framework. This allows for the extension of robust aggregators and a recent meta-aggregator to their weighted versions, mitigating the effects of delayed updates. By further incorporating a recent variance-reduction technique, we achieve an optimal convergence rate for the first time in an asynchronous Byzantine environment. Our methodology is rigorously validated through empirical and theoretical analysis, demonstrating its effectiveness in enhancing fault tolerance and optimizing performance in asynchronous ML systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de656a6f-d6f9-4287-ad46-32193481df60Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 192 citations
- Zeno++: Robust Fully Asynchronous SGDCong Xie, Sanmi Koyejo, Indranil GuptaICML 2020 · 137 citations
- Byzantine Machine Learning Made Easy By Resilient Averaging of MomentumsSadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot et al.ICML 2022 · 96 citations
- Asynchronous SGD Beats Minibatch SGD Under Arbitrary DelaysKonstantin Mishchenko, Francis R. Bach, Mathieu Even, Blake E. WoodworthNeurIPS 2022 · 95 citations
Related papers
- Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous TrainingTehila Dahan, Kfir Yehuda LevyICML 2024 · 3 citations
- Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex OptimizationHamidreza Almasi, Harsh Mishra, Balajee Vamanan, Sathya N. RaviICLR 2024
- Variance Reduction is an Antidote to Byzantines: Better Rates, Weaker Assumptions and Communication Compression as a Cherry on the TopEduard Gorbunov, Samuel Horváth, Peter Richtárik, Gauthier GidelICLR 2023 · 5 citations
- Byzantine-Tolerant Methods for Distributed Variational InequalitiesNazarii Tupitsa, Abdulla Jasem Almansoori, Yanlin Wu, Martin Takác et al.NeurIPS 2023 · 3 citations
- Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine WorkersRon Dorfman, Naseem Yehya, Kfir Yehuda LevyICML 2024 · 5 citations
