Directional Smoothness and Gradient Methods: Convergence and Adaptivity
Aaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio, Robert M. Gower
Abstract
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a measure of gradient variation that we use to develop upper-bounds on the objective. Minimizing these upper-bounds requires solving implicit equations to obtain a sequence of strongly adapted step-sizes; we show that these equations are straightforward to solve for convex quadratics and lead to new guarantees for two classical step-sizes. For general functions, we prove that the Polyak step-size and normalized GD obtain fast, path-dependent rates despite using no knowledge of the directional smoothness. Experiments on logistic regression show our convergence guarantees are tighter than the classical theory based on -smoothness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9887bc9-3e8d-4f1e-b15c-c2eb2940a4f4Cited by top-tier papers7
- New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative ResultsFrancesco Orabona, Ryan D'OrazioNeurIPS 2025 · 9 citations
- Non-Euclidean Gradient Descent Operates at the Edge of StabilityRustem Islamov, Michael Crawshaw, Jeremy Cohen, Robert GowerICML 2026 · 5 citations
- Sparse Polyak: an adaptive step size rule for high-dimensional M-estimationTianqi Qiao, Marie MarosNeurIPS 2025 · 2 citations
- Sign-SGD via Parameter-Free OptimizationDaniil Medyakov, Sergey Stanko, Gleb Molodtsov, Philip Zmushko et al.ICLR 2026 · 1 citation
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt et al.ICML 2026
Builds on11
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 171 citations
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 139 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin et al.NeurIPS 2023 · 93 citations
Related papers
- Toward a Unified Theory of Gradient Descent under Generalized SmoothnessAlexander TyurinICML 2025
- Methods for Convex (L0, L1)-Smooth Optimization: Clipping, Acceleration, and AdaptivityEduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury, Alen Aliev et al.ICLR 2025
- Optimizing (L0, L1)-Smooth Functions by Gradient MethodsDaniil Vankov, Anton Rodomanov, Angelia Nedich, Lalitha Sankar et al.ICLR 2025
- Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient DescentSharan Vaswani, Benjamin Dubois-Taine, Reza BabanezhadICML 2022
- Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive MethodsJunchi Yang, Xiang Li, Ilyas Fatkhullin, Niao HeNeurIPS 2023 · 33 citations
