Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, Suvrit Sra
Abstract
Most existing analyses of (stochastic) gradient descent rely on the condition that for -smooth costs, the step size is less than . However, many works have observed that in machine learning applications step sizes often do not fulfill this condition, yet (stochastic) gradient descent still converges, albeit in an unstable manner. We investigate this unstable convergence phenomenon from first principles, and discuss key causes behind it. We also identify its main characteristics, and how they interrelate based on both theory and experiments, offering a principled view toward understanding the phenomenon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers46
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 139 citations
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
- Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of StabilityZixuan Wang, Zhouzi Li, Jian LiNeurIPS 2022 · 71 citations
- Implicit Bias of the Step Size in Linear Diagonal Neural NetworksMor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, Daniel SoudryICML 2022 · 57 citations
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 46 citations
Builds on1
Related papers
- On the Unstable Convergence Regime of Gradient DescentShuo Chen, Jiaying Peng, Xiaolong Li, Yao ZhaoAAAI 2024 · 1 citation
- Beyond the Edge of Stability via Two-step Gradient UpdatesLei Chen, Joan BrunaICML 2023 · 22 citations
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt et al.ICML 2026
- Special Properties of Gradient Descent with Large Learning RatesAmirkeivan Mohtashami, Martin Jaggi, Sebastian U. StichICML 2023 · 16 citations
- Constant Stepsize Local GD for Logistic Regression: Acceleration by InstabilityMichael Crawshaw, Blake Woodworth, Mingrui LiuICML 2025
