ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!
Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter Richtárik
Abstract
We introduce ProxSkip -- a surprisingly simple and provably efficient method for minimizing the sum of a smooth () and an expensive nonsmooth proximable () function. The canonical approach to solving such problems is via the proximal gradient descent (ProxGD) algorithm, which is based on the evaluation of the gradient of and the prox operator of in each iteration. In this work we are specifically interested in the regime in which the evaluation of prox is costly relative to the evaluation of the gradient, which is the case in many applications. ProxSkip allows for the expensive prox operator to be skipped in most iterations: while its iteration complexity is , where is the condition number of , the number of prox evaluations is only. Our main motivation comes from federated learning, where evaluation of the gradient operator corresponds to taking a local GD step independently on all devices, and evaluation of prox corresponds to (expensive) communication in the form of gradient averaging. In this context, ProxSkip offers an effective acceleration of communication complexity. Unlike other local gradient-type methods, such as FedAvg, SCAFFOLD, S-Local-GD and FedLin, whose theoretical communication complexity is worse than, or at best matching, that of vanilla GD in the heterogeneous data regime, we obtain a provable and large improvement without any heterogeneity-bounding assumptions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06107fad-ce79-4512-9cff-dbd6324e06e6Cited by top-tier papers61
- CocktailSGD: Fine-tuning Foundation Models over 500Mbps NetworksJue Wang, Yucheng Lu, Binhang Yuan, Beidi Chen et al.ICML 2023 · 60 citations
- Momentum Provably Improves Error Feedback!Ilyas Fatkhullin, Alexander Tyurin, Peter RichtárikNeurIPS 2023 · 47 citations
- Adaptive SGD with Polyak stepsize and Line-search: Robust Convergence and Variance ReductionXiaowen Jiang, Sebastian U. StichNeurIPS 2023 · 40 citations
- TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent KernelsYaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma et al.NeurIPS 2022 · 38 citations
- Hierarchical Federated Learning with Multi-Timescale Gradient CorrectionWenzhi Fang, Dong-Jun Han, Evan Chen, Shiqiang Wang et al.NeurIPS 2024 · 34 citations
Builds on16
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 462 citations
- Is Local SGD Better than Minibatch SGD?Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai et al.ICML 2020 · 277 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
Related papers
- Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact ProxAbdurakhmon Sadiev, Dmitry Kovalev, Peter RichtárikNeurIPS 2022 · 1 citation
- Tighter Analysis for ProxSkipZhengmian Hu, Heng HuangICML 2023 · 11 citations
- Proximal and Federated Random ReshufflingKonstantin Mishchenko, Ahmed Khaled, Peter RichtárikICML 2022 · 39 citations
- Variance Reduced ProxSkip: Algorithm, Theory and Application to Federated LearningGrigory Malinovsky, Kai Yi, Peter RichtárikNeurIPS 2022 · 52 citations
- EFSkip: A New Error Feedback with Linear Speedup for Compressed Federated Learning with Arbitrary Data HeterogeneityHongyan Bao, Pengwen Chen, Ying Sun, Zhize LiAAAI 2025 · 6 citations
