FedHyper: A Universal and Robust Learning Rate Scheduler for Federated Learning with Hypergradient Descent
Ziyao Wang, Jianyu Wang, Ang Li
Abstract
The theoretical landscape of federated learning (FL) undergoes rapid evolution, but its practical application encounters a series of intricate challenges, and hyperparameter optimization is one of these critical challenges. Amongst the diverse adjustments in hyperparameters, the adaptation of the learning rate emerges as a crucial component, holding the promise of significantly enhancing the efficacy of FL systems. In response to this critical need, this paper presents FedHyper, a novel hypergradient-based learning rate adaptation algorithm specifically designed for FL. FedHyper serves as a universal learning rate scheduler that can adapt both global and local rates as the training progresses. In addition, FedHyper not only showcases unparalleled robustness to a spectrum of initial learning rate configurations but also significantly alleviates the necessity for laborious empirical learning rate adjustments. We provide a comprehensive theoretical analysis of FedHyper's convergence rate and conduct extensive experiments on vision and language benchmark datasets. The results demonstrate that FEDHYPER consistently converges 1.1-3x faster than FedAvg and the competing baselines while achieving superior final accuracy. Moreover, FedHyper catalyzes a remarkable surge in accuracy, augmenting it by up to 15% compared to FedAvg under suboptimal initial learning rate settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa78242d-3e90-40f8-bcf6-a8aba8f17218Cited by top-tier papers3
- FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank AdaptationsZiyao Wang, Zheyu Shen, Yexiao He, Guoheng Sun et al.NeurIPS 2024 · 227 citations
- DeepAFL: Deep Analytic Federated LearningJianheng Tang, Yajiang Huang, Kejia Fan, Feijiang Han et al.ICLR 2026 · 5 citations
- Provable and Practical Online Learning Rate Adaptation with Hypergradient DescentYa-Chi Chu, Wenzhi Gao, Yinyu Ye, Madeleine UdellICML 2025
Builds on6
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett et al.ICLR 2021 · 1,917 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Seizing Critical Learning Periods in Federated LearningGang Yan, Hao Wang, Jian LiAAAI 2022 · 51 citations
Related papers
- FedAdamom: Adaptive Momentum for Improved Generalization in Federated OptimizationWenjie Hou, Tianxiang Chen, Feng Wang, Tiantong Wu et al.CVPR 2026
- On the Role of Server Momentum in Federated LearningJianhui Sun, Xidong Wu, Heng Huang, Aidong ZhangAAAI 2024
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li et al.NeurIPS 2021 · 111 citations
- Problem-Parameter-Free Federated LearningWenjing Yan, Kai Zhang, Xiaolu Wang, Xuanyu CaoICLR 2025
- FedNAR: Federated Optimization with Normalized Annealing RegularizationJunbo Li, Ang Li, Chong Tian, Qirong Ho et al.NeurIPS 2023 · 11 citations
