Byzantine Machine Learning Made Easy By Resilient Averaging of Momentums
Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot, John Stephan
Abstract
Byzantine resilience emerged as a prominent topic within the distributed machine learning community. Essentially, the goal is to enhance distributed optimization algorithms, such as distributed SGD, in a way that guarantees convergence despite the presence of some misbehaving (a.k.a., Byzantine) workers. Although a myriad of techniques addressing the problem have been proposed, the field arguably rests on fragile foundations. These techniques are hard to prove correct and rely on assumptions that are (a) quite unrealistic, i.e., often violated in practice, and (b) heterogeneous, i.e., making it difficult to compare approaches. We present RESAM (RESilient Averaging of Momentums), a unified framework that makes it simple to establish optimal Byzantine resilience, relying only on standard machine learning assumptions. Our framework is mainly composed of two operators: resilient averaging at the server and distributed momentum at the workers. We prove a general theorem stating the convergence of distributed SGD under RESAM. Interestingly, demonstrating and comparing the convergence of many existing techniques become direct corollaries of our theorem, without resorting to stringent assumptions. We also present an empirical evaluation of the practical relevance of RESAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Robust Distributed Learning: Tight Error Bounds and Breakdown Point under Data HeterogeneityYoussef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot et al.NeurIPS 2023 · 37 citations
- On the Privacy-Robustness-Utility Trilemma in Distributed LearningYoussef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot et al.ICML 2023 · 33 citations
- Byzantine-Robust Learning on Heterogeneous Data via Gradient SplittingYuchen Liu, Chen Chen, Lingjuan Lyu, Fangzhao Wu et al.ICML 2023 · 27 citations
- Robust Collaborative Learning with Linear Gradient OverheadSadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang et al.ICML 2023 · 24 citations
- BFTBrain: Adaptive BFT Consensus with Reinforcement LearningChenyuan Wu, Haoyun Qin, Mohammad Javad Amiri, Boon Thau Loo et al.NSDI 2025 · 17 citations
Builds on6
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 192 citations
- Collaborative Learning in the Jungle (Decentralized, Byzantine, Heterogeneous, Asynchronous and Nonconvex Learning)El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Arsany Guirguis et al.NeurIPS 2021 · 114 citations
- Distributed Momentum for Byzantine-resilient Stochastic Gradient DescentEl Mahdi El Mhamdi, Rachid Guerraoui, Sébastien RouaultICLR 2021 · 71 citations
- Byzantine-Resilient High-Dimensional SGD with Local Iterations on Heterogeneous DataDeepesh Data, Suhas N. DiggaviICML 2021 · 49 citations
Related papers
- On the Effect of Batch Size in Byzantine-Robust Distributed LearningYi-Rui Yang, Chang-Wei Shi, Wu-Jun LiICLR 2024 · 4 citations
- Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous TrainingTehila Dahan, Kfir Yehuda LevyICML 2024 · 3 citations
- Single-Loop Byzantine-Resilient Federated Bilevel OptimizationYangnan Li, Shenghui Song, Xuanyu CaoICLR 2026
- Improving the Robustness-Utility Trade-off in Decentralized Learning over Sparse NetworksYangnan Li, Xuanyu Cao, Shenghui SongICML 2026
- Simple Minimax Optimal Byzantine Robust Algorithm for Nonconvex Objectives with Uniform Gradient HeterogeneityTomoya Murata, Kenta Niwa, Takumi Fukami, Iifan TyouICLR 2024
