SWIFT: Rapid Decentralized Federated Learning via Wait-Free Model Communication
Marco Bornstein, Tahseen Rabbani, Evan Z. Wang, Amrit S. Bedi, Furong Huang
Abstract
The decentralized Federated Learning (FL) setting avoids the role of a potentially unreliable or untrustworthy central host by utilizing groups of clients to collaboratively train a model via localized training and model/gradient sharing. Most existing decentralized FL algorithms require synchronization of client models where the speed of synchronization depends upon the slowest client. In this work, we propose SWIFT: a novel wait-free decentralized FL algorithm that allows clients to conduct training at their own speed. Theoretically, we prove that SWIFT matches the gold-standard iteration convergence rate O(1/ √ T ) of parallel stochastic gradient descent for convex and non-convex smooth optimization (total iterations T ). Furthermore, we provide theoretical results for IID and non-IID settings without any bounded-delay assumption for slow clients which is required by other asynchronous decentralized FL algorithms. Although SWIFT achieves the same iteration convergence rate with respect to T as other state-of-the-art (SOTA) parallel stochastic algorithms, it converges faster with respect to run-time due to its wait-free structure. Our experimental results demonstrate that SWIFT's run-time is reduced due to a large reduction in communication time per epoch, which falls by an order of magnitude compared to synchronous counterparts. Furthermore, SWIFT produces loss levels for image classification, over IID and non-IID data settings, upwards of 50% faster than existing SOTA algorithms. INTRODUCTION Federated Learning (FL) is an increasingly popular setting to train powerful deep neural networks with data derived from an assortment of clients. Recent research (Lian et al., 2017; Li et al., 2019; Wang & Joshi, 2018) has focused on constructing decentralized FL algorithms that overcome speed and scalability issues found within classical centralized FL (McMahan et al., 2017; Savazzi et al., 2020) . While decentralized algorithms have eliminated a major bottleneck in the distributed setting, the central server, their scalability potential is still largely untapped. Many are plagued by high communication time per round (Wang et al., 2019) . Shortening the communication time per round allows more clients to connect and then communicate with one another, thereby increasing scalability. Due to the synchronous nature of current decentralized FL algorithms, communication time per round, and consequently run-time, is amplified by parallelization delays. These delays are caused by the slowest client in the network. To circumvent these issues, asynchronous decentralized FL algorithms have been proposed (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Spectral Co-Distillation for Personalized Federated LearningZihan Chen, Howard H. Yang, Tony Q. S. Quek, Kai Fong Ernest ChongNeurIPS 2023 · 30 citations
- Tackling the Data Heterogeneity in Asynchronous Federated Learning with Cached Update CalibrationYujia Wang, Yuanpu Cao, Jingcheng Wu, Ruoyu Chen et al.ICLR 2024 · 26 citations
- Decentralized SGD and Average-direction SAM are Asymptotically EquivalentTongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song et al.ICML 2023 · 21 citations
- FADAS: Towards Federated Adaptive Asynchronous OptimizationYujia Wang, Shiqiang Wang, Songtao Lu, Jinghui ChenICML 2024 · 14 citations
- Boosting Asynchronous Decentralized Learning with Model FragmentationSayan Biswas, Anne-Marie Kermarrec, Alexis Marouani, Rafael Pires et al.WWW 2025 · 6 citations
Builds on3
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized TrainingQinyi Luo, Jiaao He, Youwei Zhuo, Xuehai QianASPLOS 2020 · 73 citations
- Asynchronous Decentralized SGD with Quantized and Local UpdatesGiorgi Nadiradze, Amirmojtaba Sabour, Peter Davies, Shigang Li et al.NeurIPS 2021 · 61 citations
Related papers
- Decentralized Sporadic Federated Learning: A Unified Algorithmic Framework with Convergence GuaranteesShahryar Zehtabi, Dong-Jun Han, Rohit Parasnis, Seyyedali Hosseinalipour et al.ICLR 2025
- SwiftFL: Enabling Speculative Training for On-Device Federated Deep LearningYuhui Zhang, Guang Yan, Xin Zhang, Zimu Guo et al.EuroSys 2026
- Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order OptimizationZhe Li, Bicheng Ying, Zidong Liu, Chaosheng Dong et al.ICLR 2025
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 95 citations
- HADFL: Heterogeneity-aware Decentralized Federated Learning FrameworkJing Cao, Zirui Lian, Weihong Liu, Zongwei Zhu et al.DAC 2021 · 28 citations
