Variance Reduction is an Antidote to Byzantines: Better Rates, Weaker Assumptions and Communication Compression as a Cherry on the Top
Eduard Gorbunov, Samuel Horváth, Peter Richtárik, Gauthier Gidel
Abstract
Byzantine-robustness has been gaining a lot of attention due to the growth of the interest in collaborative and federated learning. However, many fruitful directions, such as the usage of variance reduction for achieving robustness and communication compression for reducing communication costs, remain weakly explored in the field. This work addresses this gap and proposes Byz-VR-MARINA - a new Byzantine-tolerant method with variance reduction and compression. A key message of our paper is that variance reduction is key to fighting Byzantine workers more effectively. At the same time, communication compression is a bonus that makes the process more communication efficient. We derive theoretical convergence guarantees for Byz-VR-MARINA outperforming previous state-of-the-art for general non-convex and Polyak-Lojasiewicz loss functions. Unlike the concurrent Byzantine-robust methods with variance reduction and/or compression, our complexity results are tight and do not rely on restrictive assumptions such as boundedness of the gradients or limited compression. Moreover, we provide the first analysis of a Byzantine-tolerant method supporting non-uniform sampling of stochastic gradients. Numerical experiments corroborate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 569e16da-c7e8-46b0-914c-379a5ce869ccCited by top-tier papers6
- Robust Distributed Learning: Tight Error Bounds and Breakdown Point under Data HeterogeneityYoussef Allouah, Rachid Guerraoui, Nirupam Gupta, Rafael Pinot et al.NeurIPS 2023 · 37 citations
- Byzantine-Robust Federated Learning: Impact of Client Subsampling and Local UpdatesYoussef Allouah, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta et al.ICML 2024 · 16 citations
- On the Effect of Batch Size in Byzantine-Robust Distributed LearningYi-Rui Yang, Chang-Wei Shi, Wu-Jun LiICLR 2024 · 4 citations
- Tight Stability Bounds for Robust Distributed Learning: Byzantine Failures Hurt Generalization More than Data PoisoningThomas Boudou, Batiste Le Bars, Nirupam Gupta, Aurélien BelletICML 2026 · 3 citations
- Detecting Backdoor Attacks in Federated Learning via Direction Alignment InspectionJiahao Xu, Zikai Zhang, Rui HuCVPR 2025
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Learning from History for Byzantine Robust OptimizationSai Praneeth Karimireddy, Lie He, Martin JaggiICML 2021 · 247 citations
- EF21: A New, Simpler, Theoretically Better, and Practically Faster Error FeedbackPeter Richtárik, Igor Sokolov, Ilyas FatkhullinNeurIPS 2021 · 219 citations
- Byzantine-Robust Learning on Heterogeneous Datasets via BucketingSai Praneeth Karimireddy, Lie He, Martin JaggiICLR 2022 · 192 citations
Related papers
- Byzantine-Tolerant Methods for Distributed Variational InequalitiesNazarii Tupitsa, Abdulla Jasem Almansoori, Yanlin Wu, Martin Takác et al.NeurIPS 2023 · 3 citations
- Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous MLTehila Dahan, Kfir Y. LevyNeurIPS 2024 · 4 citations
- Secure Distributed Training at ScaleEduard Gorbunov, Alexander Borzunov, Michael Diskin, Max RyabininICML 2022 · 18 citations
- Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient DifferencesGrigory Malinovsky, Peter Richtárik, Samuel Horváth, Eduard GorbunovNeurIPS 2024 · 7 citations
- Byzantine-Resilient Non-Convex Stochastic Gradient DescentZeyuan Allen-Zhu, Faeze Ebrahimianghazani, Jerry Li, Dan AlistarhICLR 2021 · 18 citations
