Efficient Adaptive Federated Optimization
Su Hyeong Lee, Sidharth Sharma, Manzil Zaheer, Tian Li
Abstract
Adaptive optimization is critical in federated learning, where enabling adaptivity on both the server and client sides has proven essential for achieving optimal performance. However, the scalability of such jointly adaptive systems is often hindered by resource limitations in communication and memory. In this paper, we introduce a class of efficient adaptive algorithms, named FedAda 2 and its enhanced version FedAda 2 ++, designed specifically for large-scale, cross-device federated environments. FedAda 2 optimizes communication efficiency by avoiding the transfer of preconditioners between the server and clients. Additionally, FedAda 2 ++ extends this approach by incorporating memory-efficient adaptive optimizers on the client side, further reducing on-device memory usage. Theoretically, we demonstrate that FedAda 2 and FedAda 2 ++ achieve the same convergence rates for general, non-convex objectives as its more resource-intensive counterparts that directly integrate joint adaptivity. Extensive empirical evaluations on image and text datasets demonstrate both the advantages of joint adaptivity and the effectiveness and efficiency of FedAda 2 /FedAda 2 ++.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3ba367d-21ed-4db1-8feb-b32ff7d7fa5cCited by top-tier papers2
- Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal ConvergenceShuhua Yu, Dusan Jakovetic, Soummya KarICLR 2026 · 7 citations
- Expected Returns and Policy Inconsistency-Aware Offline Federated Deep Reinforcement LearningMeng XU, Zhongying Chen, Weiwei Fu, Yan Li et al.ICML 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
Related papers
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Communication-Efficient Adaptive Federated LearningYujia Wang, Lu Lin, Jinghui ChenICML 2022 · 101 citations
- Faster Adaptive Federated LearningXidong Wu, Feihu Huang, Zhengmian Hu, Heng HuangAAAI 2023 · 99 citations
- Theoretical Convergence Guaranteed Resource-Adaptive Federated Learning with Mixed HeterogeneityYangyang Wang, Xiao Zhang, Mingyi Li, Tian Lan et al.KDD 2023 · 23 citations
- DAdaQuant: Doubly-adaptive quantization for communication-efficient Federated LearningRobert Hönig, Yiren Zhao, Robert MullinsICML 2022 · 87 citations
