Adaptive Proximal Gradient Method for Convex Optimization
Yura Malitsky, Konstantin Mishchenko
Abstract
In this paper, we explore two fundamental first-order algorithms in convex optimization, namely, gradient descent (GD) and proximal gradient method (ProxGD). Our focus is on making these algorithms entirely adaptive by leveraging local curvature information of smooth functions. We propose adaptive versions of GD and ProxGD that are based on observed gradient differences and, thus, have no added computational costs. Moreover, we prove convergence of our methods assuming only local Lipschitzness of the gradient. In addition, the proposed versions allow for even larger stepsizes than those initially suggested in [MM20].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Adaptive Proximal Gradient Methods Are Universal Without ApproximationKonstantinos A. Oikonomidis, Emanuel Laude, Puya Latafat, Andreas Themelis et al.ICML 2024 · 13 citations
- Achieving Linear Convergence with Parameter-Free Algorithms in Decentralized OptimizationIlya A. Kuruzov, Gesualdo Scutari, Alexander V. GasnikovNeurIPS 2024 · 10 citations
- Universal Gradient Methods for Stochastic Convex OptimizationAnton Rodomanov, Ali Kavis, Yongtao Wu, Kimon Antonakopoulos et al.ICML 2024 · 8 citations
- Stochastic Weakly Convex Optimization beyond Lipschitz ContinuityWenzhi Gao, Qi DengICML 2024 · 6 citations
- Acceleration via silver step-size on Riemannian manifolds with applications to Wasserstein spaceJiyoung Park, Abhishek Roy, Jonathan W. Siegel, Anirban BhattacharyaNeurIPS 2025 · 3 citations
Builds on5
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 171 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
- DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent MethodAhmed Khaled, Konstantin Mishchenko, Chi JinNeurIPS 2023 · 49 citations
- A first-order primal-dual method with adaptivity to local smoothnessMaria-Luiza Vladarean, Yura Malitsky, Volkan CevherNeurIPS 2021 · 24 citations
Related papers
- Local Curvature Descent: Squeezing More Curvature out of Standard and Polyak Gradient DescentPeter Richtárik, Simone Maria Giancola, Dymitr Lubczyk, Robin YadavNeurIPS 2025 · 2 citations
- Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex OptimizationEkaterina Borodich, Dmitry KovalevICLR 2026 · 10 citations
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order InformationMajid Jahani, Sergey Rusakov, Zheng Shi, Peter Richtárik et al.ICLR 2022 · 31 citations
- Methods for Convex (L0, L1)-Smooth Optimization: Clipping, Acceleration, and AdaptivityEduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury, Alen Aliev et al.ICLR 2025
- Convergence of Clipped SGD on Convex (L0, L1)-Smooth FunctionsOfir Gaash, Kfir Y. Levy, Yair CarmonNeurIPS 2025 · 5 citations
