GradMA: A Gradient-Memory-based Accelerated Federated Learning with Alleviated Catastrophic Forgetting
Kangyang Luo, Xiang Li, Yunshi Lan, Ming Gao
Abstract
Federated Learning (FL) has emerged as a de facto machine learning area and received rapid increasing research interests from the community. However, catastrophic forgetting caused by data heterogeneity and partial participation poses distinctive challenges for FL, which are detrimental to the performance. To tackle the problems, we propose a new FL approach (namely GradMA), which takes inspiration from continual learning to simultaneously correct the server-side and worker-side update directions as well as take full advantage of server's rich computing and memory resources. Furthermore, we elaborate a memory reduction strategy to enable GradMA to accommodate FL with a large scale of workers. We then analyze convergence of GradMA theoretically under the smooth non-convex setting and show that its convergence rate achieves a linear speed up w.r.t the increasing number of sampled active workers. At last, our extensive experiments on various image classification tasks show that GradMA achieves significant performance gains in accuracy and communication efficiency compared to SOTA baselines. We provide our code here: https://github.com/lkyddd/GradMA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed InferencesMengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li et al.USENIX ATC 2024 · 78 citations
- FedAS: Bridging Inconsistency in Personalized Federated LearningXiyuan Yang, Wenke Huang, Mang YeCVPR 2024 · 69 citations
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li et al.NeurIPS 2023 · 64 citations
- FedNLR: Federated Learning with Neuron-wise Learning RatesHaozhao Wang, Peirong Zheng, Xingshuo Han, Wenchao Xu et al.KDD 2024 · 15 citations
- FedCFA: Alleviating Simpson's Paradox in Model Aggregation with Counterfactual Federated LearningZhonghua Jiang, Jimin Xu, Shengyu Zhang, Tao Shen et al.AAAI 2025 · 11 citations
Builds on24
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 957 citations
Related papers
- Data Heterogeneity and Forgotten Labels in Split Federated LearningJoana Tirana, Dimitra Tsigkari, David Solans Noguero, Nicolas KourtellisAAAI 2026 · 3 citations
- Communication-Efficient Federated Learning with Accelerated Client GradientGeeho Kim, Jinkyu Kim, Bohyung HanCVPR 2024
- Learn from Others and Be Yourself in Heterogeneous Federated LearningWenke Huang, Mang Ye, Bo DuCVPR 2022 · 254 citations
- FedAGC: Federated Continual Learning with Asymmetric Gradient CorrectionChengchao Zhang, Fanhua Shang, Hongyin Liu, Liang Wan et al.ICCV 2025 · 3 citations
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated LearningHaibo Yang, Minghong Fang, Jia LiuICLR 2021 · 310 citations
