Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning Better
Sameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan Chen
Abstract
Federated learning (FL) enables distribution of machine learning workloads from the cloud to resource-limited edge devices. Unfortunately, current deep networks remain not only too compute-heavy for inference and training on edge devices, but also too large for communicating updates over bandwidth-constrained networks. In this paper, we develop, implement, and experimentally validate a novel FL framework termed Federated Dynamic Sparse Training (FedDST) by which complex neural networks can be deployed and trained with substantially improved efficiency in both on-device computation and in-network communication. At the core of FedDST is a dynamic process that extracts and trains sparse sub-networks from the target full network. With this scheme, "two birds are killed with one stone:'' instead of full models, each client performs efficient training of its own sparse networks, and only sparse networks are transmitted between devices and the cloud. Furthermore, our results reveal that the dynamic sparsity during FL training more flexibly accommodates local heterogeneity in FL agents than the fixed, shared sparse masks. Moreover, dynamic sparsity naturally introduces an "in-time self-ensembling effect'' into the training dynamics, and improves the FL performance even over dense training. In a realistic and challenging non i.i.d. FL setting, FedDST consistently outperforms competing algorithms in our experiments: for instance, at any fixed upload data cap on non-iid CIFAR-10, it gains an impressive accuracy advantage of 10% over FedAvgM when given the same upload data cap; the accuracy gap remains 3% even when FedAvgM is given 2 times the upload data cap, further demonstrating efficacy of FedDST. Code is available at: https://github.com/bibikar/feddst.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eabeba28-5302-45fe-b17a-a85832076095Cited by top-tier papers30
- DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse TrainingRong Dai, Li Shen, Fengxiang He, Xinmei Tian et al.ICML 2022 · 163 citations
- Efficient Personalized Federated Learning via Sparse Model-AdaptationDaoyuan Chen, Liuyi Yao, Dawei Gao, Bolin Ding et al.ICML 2023 · 76 citations
- Resource-Adaptive Federated Learning with All-In-One Neural CompositionYiqun Mei, Pengfei Guo, Mo Zhou, Vishal PatelNeurIPS 2022 · 62 citations
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic SparsityShiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen et al.ICLR 2022 · 62 citations
- FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel ExtractionFeijie Wu, Xingchen Wang, Yaqing Wang, Tianci Liu et al.NeurIPS 2024 · 47 citations
Builds on12
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- FedASMU: Efficient Asynchronous Federated Learning with Dynamic Staleness-Aware Model UpdateJi Liu, Juncheng Jia, Tianshi Che, Chao Huo et al.AAAI 2024 · 87 citations
- FedFit: Federated Dynamic Sparse Training via Fisher Information scoringMeng Bi, Hong Huang, Jinlong Song, Charles Wang et al.ICML 2026
- Hermes: an efficient federated learning framework for heterogeneous mobile clientsAng Li, Jingwei Sun, Pengcheng Li, Yu Pu et al.MobiCom 2021 · 167 citations
- Communication-Efficient Federated Learning for Heterogeneous Edge Devices Based on Adaptive Gradient QuantizationHeting Liu, Fang He, Guohong CaoINFOCOM 2023 · 60 citations
- ScaleFL: Resource-Adaptive Federated Learning with Heterogeneous ClientsFatih Ilhan, Gong Su, Ling LiuCVPR 2023
