Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Alexander Tyurin
Abstract
We study the classical optimization problem min x∈R d f (x) and analyze the gradient descent (GD) method in both nonconvex and convex settings. It is well-known that, under the L-smoothness assumption (∥∇ 2 f (x)∥ ≤ L), the optimal point minimizing the quadratic upper bound f Surprisingly, a similar result can be derived under the ℓ-generalized smoothness assumption (∥∇ 2 f (x)∥ ≤ ℓ(∥∇f (x)∥)). In this case, we derive the step size . Using this step size rule, we improve upon existing theoretical convergence rates and obtain new results in several previously unexplored setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c3abb1d-c3f6-424b-89a7-40e0d62d882eCited by top-tier papers6
- Gradient-Normalized Smoothness for Optimization with Approximate HessiansAndrei Semenov, Martin Jaggi, Nikita DoikovICLR 2026 · 8 citations
- Convergence of Clipped SGD on Convex (L0, L1)-Smooth FunctionsOfir Gaash, Kfir Y. Levy, Yair CarmonNeurIPS 2025 · 5 citations
- On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE ApproachEnea Monzio Compagnoni, Rustem Islamov, Frank Proske, Aurelien Lucchi et al.ICML 2026 · 4 citations
- Near-Optimal Convergence of Accelerated Gradient Methods under Generalized and -SmoothnessAlexander TyurinICML 2026 · 1 citation
- Taming Stochastic Gradient Descent: Almost Sure Convergence and Saddle-Point Avoidance under -SmoothnessVassilis Apidopoulos, Iosif Lytras, Panayotis MertikopoulosICML 2026
Builds on8
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 132 citations
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang et al.NeurIPS 2022 · 111 citations
- Revisiting Gradient Clipping: Stochastic bias and tight convergence guaranteesAnastasia Koloskova, Hadrien Hendrikx, Sebastian U. StichICML 2023 · 106 citations
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin et al.NeurIPS 2023 · 93 citations
Related papers
- Methods for Convex (L0, L1)-Smooth Optimization: Clipping, Acceleration, and AdaptivityEduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury, Alen Aliev et al.ICLR 2025
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio et al.NeurIPS 2024 · 25 citations
- Optimizing (L0, L1)-Smooth Functions by Gradient MethodsDaniil Vankov, Anton Rodomanov, Angelia Nedich, Lalitha Sankar et al.ICLR 2025
- The Complexity of Finding Stationary Points with Stochastic Gradient DescentYoel Drori, Ohad ShamirICML 2020 · 73 citations
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 171 citations
