Tighter Analysis for ProxSkip
Zhengmian Hu, Heng Huang
摘要
In this paper, we provide a tighter analysis for ProxSkip, an algorithm that allows fewer proximal operator computations to solve composite optimization problems. We improve the existing decreasing speed of Lyapunov function from O(p 2 ) to O(p), when p, the frequency of the proximal operators is small enough. Our theoretical analysis also reveals the drawbacks of using large step sizes for gradient descent in ProxSkip when the proximal operator part is the bottleneck. Our main motivation comes from the continuous limit in which the original analysis of ProxSkip fails to guarantee convergence when both the step size γ and frequency p tend to zero. We construct a counterexample to demonstrate why such counterintuitive behavior occurs for the original analysis and then propose a novel Lyapunov function variant to construct a tighter analysis, avoiding the problem of the old one. Such a new Lyapunov function can be directly extended to many other variants of ProxSkip. When applied to stochastic gradient setup, our analysis leads to an improved proximal operator complexity for SProxSkip from O( 1 /εµ 2 log( 1 /ε)) to O( √ κ log( 1 /ε)).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Federated Optimization with Doubly Regularized Drift CorrectionXiaowen Jiang, Anton Rodomanov, Sebastian U. StichICML 2024 · 被引用 18 次
- Stabilized Proximal-Point Methods for Federated OptimizationXiaowen Jiang, Anton Rodomanov, Sebastian U. StichNeurIPS 2024 · 被引用 13 次
- SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD LearningPaul Mangold, Sergey Samsonov, Safwan Labbi, Ilya Levin 等NeurIPS 2024 · 被引用 10 次
- Addressing Label Shift in Distributed Learning via Entropy RegularizationZhiyuan Wu, Changkyu Choi, Xiangcheng Cao, Volkan Cevher 等ICLR 2025
- Scaffold with Stochastic Gradients: New Analysis with Linear Speed-UpPaul Mangold, Alain Oliviero Durmus, Aymeric Dieuleveut, Eric MoulinesICML 2025
它引用的顶会 Paper8
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally!Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter RichtárikICML 2022 · 被引用 200 次
- Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse GradientsAritra Mitra, Rayana H. Jaafar, George J. Pappas, Hamed HassaniNeurIPS 2021 · 被引用 193 次
相关 Paper
- Non-convex Stochastic Composite Optimization with Polyak MomentumYuan Gao, Anton Rodomanov, Sebastian U. StichICML 2024 · 被引用 13 次
- Communication Acceleration of Local Gradient Methods via an Accelerated Primal-Dual Algorithm with an Inexact ProxAbdurakhmon Sadiev, Dmitry Kovalev, Peter RichtárikNeurIPS 2022 · 被引用 1 次
- EFSkip: A New Error Feedback with Linear Speedup for Compressed Federated Learning with Arbitrary Data HeterogeneityHongyan Bao, Pengwen Chen, Ying Sun, Zhize LiAAAI 2025 · 被引用 6 次
- Tighter Performance Theory of FedExProxWojciech Anyszka, Kaja Gruntkowska, Alexander Tyurin, Peter RichtárikICLR 2026 · 被引用 3 次
- Decentralized Accelerated Proximal Gradient DescentHaishan Ye, Ziang Zhou, Luo Luo, Tong ZhangNeurIPS 2020 · 被引用 37 次
