Toward a Unified Theory of Gradient Descent under Generalized Smoothness
Alexander Tyurin
2025年份
6顶会引用
摘要
We study the classical optimization problem min x∈R d f (x) and analyze the gradient descent (GD) method in both nonconvex and convex settings. It is well-known that, under the L-smoothness assumption (∥∇ 2 f (x)∥ ≤ L), the optimal point minimizing the quadratic upper bound f Surprisingly, a similar result can be derived under the ℓ-generalized smoothness assumption (∥∇ 2 f (x)∥ ≤ ℓ(∥∇f (x)∥)). In this case, we derive the step size . Using this step size rule, we improve upon existing theoretical convergence rates and obtain new results in several previously unexplored setups.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Gradient-Normalized Smoothness for Optimization with Approximate HessiansAndrei Semenov, Martin Jaggi, Nikita DoikovICLR 2026 · 被引用 8 次
- Convergence of Clipped SGD on Convex (L0, L1)-Smooth FunctionsOfir Gaash, Kfir Y. Levy, Yair CarmonNeurIPS 2025 · 被引用 5 次
- On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE ApproachEnea Monzio Compagnoni, Rustem Islamov, Frank Proske, Aurelien Lucchi 等ICML 2026 · 被引用 4 次
- Near-Optimal Convergence of Accelerated Gradient Methods under Generalized and -SmoothnessAlexander TyurinICML 2026 · 被引用 1 次
- Taming Stochastic Gradient Descent: Almost Sure Convergence and Saddle-Point Avoidance under -SmoothnessVassilis Apidopoulos, Iosif Lytras, Panayotis MertikopoulosICML 2026
它引用的顶会 Paper8
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Convergence of Adam Under Relaxed AssumptionsHaochuan Li, Alexander Rakhlin, Ali JadbabaieNeurIPS 2023 · 被引用 132 次
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang 等NeurIPS 2022 · 被引用 111 次
- Revisiting Gradient Clipping: Stochastic bias and tight convergence guaranteesAnastasia Koloskova, Hadrien Hendrikx, Sebastian U. StichICML 2023 · 被引用 106 次
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin 等NeurIPS 2023 · 被引用 93 次
相关 Paper
- Methods for Convex (L0, L1)-Smooth Optimization: Clipping, Acceleration, and AdaptivityEduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury, Alen Aliev 等ICLR 2025
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio 等NeurIPS 2024 · 被引用 25 次
- Optimizing (L0, L1)-Smooth Functions by Gradient MethodsDaniil Vankov, Anton Rodomanov, Angelia Nedich, Lalitha Sankar 等ICLR 2025
- The Complexity of Finding Stationary Points with Stochastic Gradient DescentYoel Drori, Ohad ShamirICML 2020 · 被引用 73 次
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 被引用 171 次
