Simple Minimax Optimal Byzantine Robust Algorithm for Nonconvex Objectives with Uniform Gradient Heterogeneity
Tomoya Murata, Kenta Niwa, Takumi Fukami, Iifan Tyou
Abstract
In this study, we consider nonconvex federated learning problems with the existence of Byzantine workers. We propose a new simple Byzantine robust algorithm called Momentum Screening. The algorithm is adaptive to the Byzantine fraction, i.e., all its hyperparameters do not depend on the number of Byzantine workers. We show that our method achieves the best optimization error of O(δ 2 ζ 2 max ) for nonconvex smooth local objectives satisfying ζ max -uniform gradient heterogeneity condition under δ-Byzantine fraction, which can be better than the best known error rate of O(δζ 2 mean ) for local objectives satisfying ζ mean -mean heterogeneity condition when δ ≤ (ζ mean /ζ max ) 2 . Furthermore, we derive an algorithm independent lower bound for local objectives satisfying ζ max -uniform gradient heterogeneity condition and show the minimax optimality of our proposed method on this class. In numerical experiments, we validate the superiority of our method over the existing robust aggregation algorithms and verify our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ded689e-f39b-4c2d-8bfe-9e7527871dedCited by top-tier papers1
Ask how each one uses itBuilds on8
- Ditto: Fair and Robust Federated Learning Through PersonalizationTian Li, Shengyuan Hu, Ahmad Beirami, Virginia SmithICML 2021 · 1,313 citations
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 192 citations
- Collaborative Learning in the Jungle (Decentralized, Byzantine, Heterogeneous, Asynchronous and Nonconvex Learning)El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Arsany Guirguis et al.NeurIPS 2021 · 114 citations
- Byzantine-Resilient High-Dimensional SGD with Local Iterations on Heterogeneous DataDeepesh Data, Suhas N. DiggaviICML 2021 · 49 citations
Related papers
- On the Tension between Byzantine Robustness and No-Attack Accuracy in Distributed LearningYi-Rui Yang, Chang-Wei Shi, Wu-Jun LiICML 2025
- Efficient Federated Learning against Byzantine Attacks and Data Heterogeneity via Aggregating Normalized GradientsShiyuan Zuo, Xingrun Yan, Rongfei Fan, Li Shen et al.NeurIPS 2025 · 8 citations
- Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial ParticipationKaoru Otsuka, Yuki Takezawa, Makoto YamadaICML 2026
- On the Effect of Batch Size in Byzantine-Robust Distributed LearningYi-Rui Yang, Chang-Wei Shi, Wu-Jun LiICLR 2024 · 4 citations
- Byzantine-Robust Federated Learning with Learnable Aggregation WeightsJavad Parsa, Amir Hossein Daghestani, André M. H. Teixeira, Mikael JohanssonICLR 2026 · 2 citations
