Fault Tolerant ML: Efficient Meta-Aggregation and Synchronous Training
Tehila Dahan, Kfir Yehuda Levy
Abstract
In this paper, we investigate the challenging framework of Byzantine-robust training in distributed machine learning (ML) systems, focusing on enhancing both efficiency and practicality. As distributed ML systems become integral for complex ML tasks, ensuring resilience against Byzantine failures-where workers may contribute incorrect updates due to malice or error-gains paramount importance. Our first contribution is the introduction of the Centered Trimmed Meta Aggregator (CTMA), an efficient meta-aggregator that upgrades baseline aggregators to optimal performance levels, while requiring low computational demands. Additionally, we propose harnessing a recently developed gradient estimation technique based on a double-momentum strategy within the Byzantine context. Our paper highlights its theoretical and practical advantages for Byzantine-robust training, especially in simplifying the tuning process and reducing the reliance on numerous hyperparameters. The effectiveness of this technique is supported by theoretical insights within the stochastic convex optimization (SCO) framework and corroborated by empirical evidence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af49ee55-cfa3-48c3-8fd0-a329fa6af719Cited by top-tier papers2
- Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous MLTehila Dahan, Kfir Y. LevyNeurIPS 2024 · 4 citations
- Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial ParticipationKaoru Otsuka, Yuki Takezawa, Makoto YamadaICML 2026
Builds on6
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 192 citations
- Byzantine Machine Learning Made Easy By Resilient Averaging of MomentumsSadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot et al.ICML 2022 · 96 citations
- Distributed Momentum for Byzantine-resilient Stochastic Gradient DescentEl Mahdi El Mhamdi, Rachid Guerraoui, Sébastien RouaultICLR 2021 · 71 citations
- Byzantine-Resilient Non-Convex Stochastic Gradient DescentZeyuan Allen-Zhu, Faeze Ebrahimianghazani, Jerry Li, Dan AlistarhICLR 2021 · 18 citations
Related papers
- Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex OptimizationHamidreza Almasi, Harsh Mishra, Balajee Vamanan, Sathya N. RaviICLR 2024
- Dynamic Byzantine-Robust Learning: Adapting to Switching Byzantine WorkersRon Dorfman, Naseem Yehya, Kfir Yehuda LevyICML 2024 · 5 citations
- Single-Loop Byzantine-Resilient Federated Bilevel OptimizationYangnan Li, Shenghui Song, Xuanyu CaoICLR 2026
- Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with OutliersYuhao Yi, Ronghui You, Hong Liu, Changxin Liu et al.AAAI 2024 · 7 citations
- High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine AttackersWenyu Liu, Tianqiang Huang, Pengfei Zhang, Zong Ke et al.AAAI 2026 · 10 citations
