SwiftFL: Enabling Speculative Training for On-Device Federated Deep Learning
Yuhui Zhang, Guang Yan, Xin Zhang, Zimu Guo, Lutan Zhao, Jiangfeng Cao, Dan Meng, Rui Hou
Abstract
Federated deep learning (FDL) is a promising privacy-preserving approach for training deep neural networks on distributed datasets without raw data sharing. But the classical synchronous FDL faces straggler problem: slow trainers severely impede overall efficiency. Inspired by speculative execution techniques in modern processors, this paper proposes SwiftFL, a novel and efficient speculative training system for FDL. Instead of simply waiting for slower trainer, SwiftFL proactively updates the global model with predicted gradients, enabling faster trainers to speculatively initiate the next training round. Furthermore, a gradient compensation technique is proposed to correct mispredicted training without re-training. Finally, to overcome the model-drift problem caused by fast trainers perform more local training rounds, we propose a client selection strategy. This strategy determines whether trainers should perform speculative training by striking a balance between two metrics: model drift degree and local training efficiency. In the evaluation, we compare SwiftFL with four state-of-the-art FDL systems and demonstrate that SwiftFL achieves an average speedup of 6.08× while maintaining consistent final model accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 180ef797-6de5-4f39-9eac-716a8d016d67Related papers
- SpecFL: An Efficient Speculative Federated Learning System for Tree-based Model TrainingYuhui Zhang, Lutan Zhao, Cheng Che, XiaoFeng Wang et al.HPCA 2024 · 3 citations
- SWIFT: Rapid Decentralized Federated Learning via Wait-Free Model CommunicationMarco Bornstein, Tahseen Rabbani, Evan Z. Wang, Amrit S. Bedi et al.ICLR 2023 · 3 citations
- On the Convergence of Communication-Efficient Local SGD for Federated LearningHongchang Gao, An Xu, Heng HuangAAAI 2021 · 66 citations
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh et al.PPoPP 2020 · 52 citations
- HADFL: Heterogeneity-aware Decentralized Federated Learning FrameworkJing Cao, Zirui Lian, Weihong Liu, Zongwei Zhu et al.DAC 2021 · 28 citations
