Tighter Performance Theory of FedExProx
Wojciech Anyszka, Kaja Gruntkowska, Alexander Tyurin, Peter Richtárik
摘要
We revisit FedExProx-a recently proposed distributed optimization method designed to enhance convergence properties of parallel proximal algorithms via extrapolation. In the process, we uncover a surprising flaw: its known theoretical guarantees on quadratic optimization tasks are no better than those offered by the vanilla Gradient Descent (GD) method. Motivated by this observation, we develop a novel analysis framework, establishing a tighter linear convergence rate for nonstrongly convex quadratic problems. By incorporating both computation and communication costs, we demonstrate that FedExProx can indeed provably outperform GD, in stark contrast to the original analysis. Furthermore, we consider partial participation scenarios and analyze two adaptive extrapolation strategies-based on gradient diversity and Polyak stepsizes-again significantly outperforming previous results. Moving beyond quadratics, we extend the applicability of our analysis to general functions satisfying the Polyak-Łojasiewicz condition, outperforming the previous strongly convex analysis while operating under weaker assumptions. Backed by empirical results, our findings point to a new and stronger potential of FedExProx, paving the way for further exploration of the benefits of extrapolation in federated learning. * The work of Wojciech Anyszka was conducted during a VSRP internship at KAUST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 被引用 200 次
- The Power of Extrapolation in Federated LearningHanmin Li, Kirill Acharya, Peter RichtárikNeurIPS 2024 · 被引用 16 次
相关 Paper
- Methods with Local Steps and Random Reshuffling for Generally Smooth Non-Convex Federated OptimizationYury Demidovich, Petr Ostroukhov, Grigory Malinovsky, Samuel Horváth 等ICLR 2025
- On Convergence of FedProx: Local Dissimilarity Invariant Bounds, Non-smoothness and BeyondXiaotong Yuan, Ping LiNeurIPS 2022 · 被引用 141 次
- FedExP: Speeding Up Federated Averaging via ExtrapolationDivyansh Jhunjhunwala, Shiqiang Wang, Gauri JoshiICLR 2023 · 被引用 8 次
- DASHA: Distributed Nonconvex Optimization with Communication Compression and Optimal Oracle ComplexityAlexander Tyurin, Peter RichtárikICLR 2023 · 被引用 2 次
- EFSkip: A New Error Feedback with Linear Speedup for Compressed Federated Learning with Arbitrary Data HeterogeneityHongyan Bao, Pengwen Chen, Ying Sun, Zhize LiAAAI 2025 · 被引用 6 次
