Delayed Gradient Averaging: Tolerate the Communication Latency for Federated Learning
Ligeng Zhu, Hongzhou Lin, Yao Lu, Yujun Lin, Song Han
Abstract
Federated Learning is an emerging direction in distributed machine learning that enables jointly training a model without sharing the data. Since the data is distributed across many edge devices through wireless / long-distance connections, federated learning suffers from inevitable high communication latency. However, the latency issues are undermined in the current literature [15] and existing approaches such as FedAvg [27] become less efficient when the latency increases. To overcome the problem, we propose Delayed Gradient Averaging (DGA), which delays the averaging step to improve efficiency and allows local computation in parallel to communication. We theoretically prove that DGA attains a similar convergence rate as FedAvg, and empirically show that our algorithm can tolerate high network latency without compromising accuracy. Specifically, we benchmark the training speed on various vision (CIFAR, ImageNet) and language tasks (Shakespeare), with both IID and non-IID partitions, and show DGA can bring 2.55⇥ to 4.07⇥ speedup. Moreover, we built a 16-node Raspberry Pi cluster and show that DGA can consistently speed up real-world federated learning applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e76b9fc3-69c7-4121-ba91-7c57278580dbCited by top-tier papers10
- Spectral Co-Distillation for Personalized Federated LearningZihan Chen, Howard H. Yang, Tony Q. S. Quek, Kai Fong Ernest ChongNeurIPS 2023 · 30 citations
- MAS: Towards Resource-Efficient Federated Multiple-Task LearningWeiming Zhuang, Yonggang Wen, Lingjuan Lyu, Shuai ZhangICCV 2023 · 22 citations
- No One Idles: Efficient Heterogeneous Federated Learning with Parallel Edge and Server ComputationFeilong Zhang, Xianming Liu, Shiyi Lin, Gang Wu et al.ICML 2023 · 15 citations
- Fedhca2: Towards Hetero-Client Federated Multi-Task LearningYuxiang Lu, Suizhi Huang, Yuwen Yang, Shalayiding Sirejiding et al.CVPR 2024 · 13 citations
- Workie-Talkie: Accelerating Federated Learning by Overlapping Computing and Communications via Contrastive RegularizationRui Chen, Qiyu Wan, Pavana Prakash, Lan Zhang et al.ICCV 2023 · 10 citations
Builds on6
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Fair Resource Allocation in Federated LearningTian Li, Maziar Sanjabi, Ahmad Beirami, Virginia SmithICLR 2020 · 971 citations
- FetchSGD: Communication-Efficient Federated Learning with SketchingDaniel Rothchild, Ashwinee Panda, Enayat Ullah, Nikita Ivkin et al.ICML 2020 · 425 citations
Related papers
- AOCC-FL: Federated Learning with Aligned Overlapping via Calibrated CompensationHaozhao Wang, Wenchao Xu, Yunfeng Fan, Ruixuan Li et al.INFOCOM 2023 · 8 citations
- FedADMM: A Robust Federated Deep Learning Framework with Adaptivity to System HeterogeneityYonghai Gong, Yichuan Li, Nikolaos M. FrerisICDE 2022 · 41 citations
- Hybrid Local SGD for Federated Learning with Heterogeneous CommunicationsYuanxiong Guo, Ying Sun, Rui Hu, Yanmin GongICLR 2022 · 63 citations
- Enhancing Decentralized Federated Learning for Non-IID Data on Heterogeneous DevicesMin Chen, Yang Xu, Hongli Xu, Liusheng HuangICDE 2023 · 25 citations
- FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient EstimationSrijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi et al.ICML 2025
