FetchSGD: Communication-Efficient Federated Learning with Sketching
Daniel Rothchild, Ashwinee Panda, Enayat Ullah, Nikita Ivkin, Ion Stoica, Vladimir Braverman, Joseph Gonzalez, Raman Arora
Abstract
Existing approaches to federated learning suffer from a communication bottleneck as well as convergence issues due to sparse client participation. In this paper we introduce a novel algorithm, called FetchSGD, to overcome these challenges. FetchSGD compresses model updates using a Count Sketch, and then takes advantage of the mergeability of sketches to combine model updates from many workers. A key insight in the design of FetchSGD is that, because the Count Sketch is linear, momentum and error accumulation can both be carried out within the sketch. This allows the algorithm to move momentum and error accumulation from clients to the central aggregator, overcoming the challenges of sparse client participation while still achieving high compression rates and good convergence. We prove that FetchSGD has favorable convergence guarantees, and we demonstrate its empirical effectiveness by training two residual networks and a transformer model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e17a975c-3401-4146-91ff-7194d2e1b005Cited by top-tier papers76
- FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleFan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Jiachen Liu et al.ICML 2022 · 280 citations
- Neurotoxin: Durable Backdoors in Federated LearningZhengming Zhang, Ashwinee Panda, Linyue Song, Yaoqing Yang et al.ICML 2022 · 209 citations
- Provably Secure Federated Learning against Malicious ClientsXiaoyu Cao, Jinyuan Jia, Neil Zhenqiang GongAAAI 2021 · 161 citations
- Personalized Federated Learning via Variational Bayesian InferenceXu Zhang, Yinchuan Li, Wenpeng Li, Kaiyang Guo et al.ICML 2022 · 132 citations
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 113 citations
Builds on1
Related papers
- On the Convergence of Communication-Efficient Local SGD for Federated LearningHongchang Gao, An Xu, Heng HuangAAAI 2021 · 66 citations
- Towards Faster Decentralized Stochastic Optimization with Communication CompressionRustem Islamov, Yuan Gao, Sebastian U. StichICLR 2025
- FedFetch: Faster Federated Learning with Adaptive Downstream PrefetchingQifan Yan, Andrew Liu, Shiqi He, Mathias Lécuyer et al.INFOCOM 2025 · 2 citations
- Momentum Benefits Non-iid Federated Learning Simply and ProvablyZiheng Cheng, Xinmeng Huang, Pengfei Wu, Kun YuanICLR 2024 · 40 citations
- Momentum Ensures Convergence of SIGNSGD under Weaker AssumptionsTao Sun, Qingsong Wang, Dongsheng Li, Bao WangICML 2023 · 36 citations
