Variance Reduced ProxSkip: Algorithm, Theory and Application to Federated Learning
Grigory Malinovsky, Kai Yi, Peter Richtárik
Abstract
We study distributed optimization methods based on the local training (LT) paradigm: achieving communication efficiency by performing richer local gradient-based training on the clients before parameter averaging. Looking back at the progress of the field, we identify 5 generations of LT methods: 1) heuristic, 2) homogeneous, 3) sublinear, 4) linear, and 5) accelerated. The 5 generation, initiated by the ProxSkip method of Mishchenko, Malinovsky, Stich and Richtárik (2022) and its analysis, is characterized by the first theoretical confirmation that LT is a communication acceleration mechanism. Inspired by this recent progress, we contribute to the 5 generation of LT methods by showing that it is possible to enhance them further using variance reduction. While all previous theoretical results for LT methods ignore the cost of local work altogether, and are framed purely in terms of the number of communication rounds, we show that our methods can be substantially faster in terms of the total training cost than the state-of-the-art method ProxSkip in theory and practice in the regime when local computation is sufficiently expensive. We characterize this threshold theoretically, and confirm our theoretical predictions with empirical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0882c11-9cff-432c-ba80-d0b531f4ceb5Cited by top-tier papers13
- Hierarchical Federated Learning with Multi-Timescale Gradient CorrectionWenzhi Fang, Dong-Jun Han, Evan Chen, Shiqiang Wang et al.NeurIPS 2024 · 34 citations
- FedP3: Federated Personalized and Privacy-friendly Network Pruning under Model HeterogeneityKai Yi, Nidham Gazagnadou, Peter Richtárik, Lingjuan LyuICLR 2024 · 18 citations
- Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational InequalitiesAleksandr Beznosikov, Martin Takác, Alexander V. GasnikovNeurIPS 2023 · 15 citations
- Tighter Analysis for ProxSkipZhengmian Hu, Heng HuangICML 2023 · 11 citations
- SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD LearningPaul Mangold, Sergey Samsonov, Safwan Labbi, Ilya Levin et al.NeurIPS 2024 · 10 citations
Builds on7
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Minibatch vs Local SGD for Heterogeneous Distributed LearningBlake E. Woodworth, Kumar Kshitij Patel, Nati SrebroNeurIPS 2020 · 231 citations
- Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse GradientsAritra Mitra, Rayana H. Jaafar, George J. Pappas, Hamed HassaniNeurIPS 2021 · 193 citations
- From Local SGD to Local Fixed-Point Methods for Federated LearningGrigory Malinovskiy, Dmitry Kovalev, Elnur Gasanov, Laurent Condat et al.ICML 2020 · 135 citations
Related papers
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 200 citations
- Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact ProxAbdurakhmon Sadiev, Dmitry Kovalev, Peter RichtárikNeurIPS 2022 · 1 citation
- Step-Ahead Error Feedback for Distributed Training with Compressed GradientAn Xu, Zhouyuan Huo, Heng HuangAAAI 2021 · 17 citations
- Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed OptimizationMher Safaryan, Filip Hanzely, Peter RichtárikNeurIPS 2021 · 32 citations
- Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local UpdatesSiqi Zhang, Sayantan Choudhury, Sebastian U. Stich, Nicolas LoizouICLR 2024 · 9 citations
