The Power of Extrapolation in Federated Learning
Hanmin Li, Kirill Acharya, Peter Richtárik
Abstract
We propose and study several server-extrapolation strategies for enhancing the theoretical and empirical convergence properties of the popular federated learning optimizer FedProx [Li et al., 2020]. While it has long been known that some form of extrapolation can help in the practice of FL, only a handful of works provide any theoretical guarantees. The phenomenon seems elusive, and our current theoretical understanding remains severely incomplete. In our work, we focus on smooth convex or strongly convex problems in the interpolation regime. In particular, we propose Extrapolated FedProx (FedExProx), and study three extrapolation strategies: a constant strategy (depending on various smoothness parameters and the number of participating devices), and two smoothness-adaptive strategies; one based on the notion of gradient diversity (FedExProx-GraDS), and the other one based on the stochastic Polyak stepsize (FedExProx-StoPS). Our theory is corroborated with carefully constructed numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and AccelerationAhmed Khaled, Satyen Kale, Arthur Douillard, Chi Jin et al.NeurIPS 2025 · 7 citations
- Tighter Performance Theory of FedExProxWojciech Anyszka, Kaja Gruntkowska, Alexander Tyurin, Peter RichtárikICLR 2026 · 3 citations
- Understanding the Statistical Accuracy-Communication Trade-off in Personalized Federated Learning with Minimax GuaranteesXin Yu, Zelin He, Ying Sun, Lingzhou Xue et al.ICML 2025
- S-D-RSM: Stochastic Distributed Regularized Splitting Method for Large-Scale Convex Optimization ProblemsMaoran Wang, Xingju Cai, Yongxin ChenAAAI 2026
- Convergence of Distributed Adaptive Optimization with Local UpdatesZiheng Cheng, Margalit GlasgowICLR 2025
Builds on8
- Personalized Federated Learning with Moreau EnvelopesCanh T. Dinh, Nguyen Hoang Tran, Tuan Dung NguyenNeurIPS 2020 · 1,542 citations
- On Convergence of FedProx: Local Dissimilarity Invariant Bounds, Non-smoothness and BeyondXiaotong Yuan, Ping LiNeurIPS 2022 · 141 citations
- Dynamics of SGD with Stochastic Polyak Stepsizes: Truly Adaptive Variants and Convergence to Exact SolutionAntonio Orvieto, Simon Lacoste-Julien, Nicolas LoizouNeurIPS 2022 · 57 citations
- Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed OptimizationMher Safaryan, Filip Hanzely, Peter RichtárikNeurIPS 2021 · 32 citations
- Minibatch Stochastic Approximate Proximal Point MethodsHilal Asi, Karan N. Chadha, Gary Cheng, John C. DuchiNeurIPS 2020 · 22 citations
Related papers
- FedExP: Speeding Up Federated Averaging via ExtrapolationDivyansh Jhunjhunwala, Shiqiang Wang, Gauri JoshiICLR 2023 · 8 citations
- Methods with Local Steps and Random Reshuffling for Generally Smooth Non-Convex Federated OptimizationYury Demidovich, Petr Ostroukhov, Grigory Malinovsky, Samuel Horváth et al.ICLR 2025
- Federated Accelerated Stochastic Gradient DescentHonglin Yuan, Tengyu MaNeurIPS 2020 · 217 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Fast Federated Learning in the Presence of Arbitrary Device UnavailabilityXinran Gu, Kaixuan Huang, Jingzhao Zhang, Longbo HuangNeurIPS 2021 · 142 citations
