Secure Distributed Training at Scale
Eduard Gorbunov, Alexander Borzunov, Michael Diskin, Max Ryabinin
Abstract
Some of the hardest problems in deep learning can be solved with the combined effort of many independent parties, as is the case for volunteer computing and federated learning. These setups rely on high numbers of peers to provide computational resources or train on decentralized datasets. Unfortunately, participants in such systems are not always reliable. Any single participant can jeopardize the entire training run by sending incorrect updates, whether deliberately or by mistake. Training in presence of such peers requires specialized distributed training algorithms with Byzantine tolerance. These algorithms often sacrifice efficiency by introducing redundant communication or passing all updates through a trusted server. As a result, it can be infeasible to apply such algorithms to large-scale distributed deep learning, where models can have billions of parameters. In this work, we propose a novel protocol for secure (Byzantine-tolerant) decentralized training that emphasizes communication efficiency. We rigorously analyze this protocol: in particular, we provide theoretical bounds for its resistance against Byzantine and Sybil attacks and show that it has a marginal communication overhead. To demonstrate its practical effectiveness, we conduct large-scale experiments on image classification and language modeling in presence of Byzantine attackers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cc23dfc-d718-408a-8765-8fde78e103e0Cited by top-tier papers4
- Robust Collaborative Learning with Linear Gradient OverheadSadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang et al.ICML 2023 · 24 citations
- Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient DifferencesGrigory Malinovsky, Peter Richtárik, Samuel Horváth, Eduard GorbunovNeurIPS 2024 · 7 citations
- Variance Reduction is an Antidote to Byzantines: Better Rates, Weaker Assumptions and Communication Compression as a Cherry on the TopEduard Gorbunov, Samuel Horváth, Peter Richtárik, Gauthier GidelICLR 2023 · 5 citations
- Byzantine-Tolerant Methods for Distributed Variational InequalitiesNazarii Tupitsa, Abdulla Jasem Almansoori, Yanlin Wu, Martin Takác et al.NeurIPS 2023 · 3 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
Related papers
- Byzantine-robust Decentralized Federated Learning via Dual-domain Clustering and Trust BootstrappingPeng Sun, Xinyang Liu, Zhibo Wang, Bo LiuCVPR 2024 · 21 citations
- Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex OptimizationHamidreza Almasi, Harsh Mishra, Balajee Vamanan, Sathya N. RaviICLR 2024
- Election Coding for Distributed Learning: Protecting SignSGD against Byzantine AttacksJy-yong Sohn, Dong-Jun Han, Beongjun Choi, Jaekyun MoonNeurIPS 2020 · 41 citations
- Efficient Federated Learning against Byzantine Attacks and Data Heterogeneity via Aggregating Normalized GradientsShiyuan Zuo, Xingrun Yan, Rongfei Fan, Li Shen et al.NeurIPS 2025 · 8 citations
- Byzantine-Robust Decentralized Federated LearningMinghong Fang, Zifan Zhang, Hairi, Prashant Khanduri et al.CCS 2024 · 38 citations
