Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, Suvrit Sra
2022年份
89被引次数
46顶会引用
摘要
Most existing analyses of (stochastic) gradient descent rely on the condition that for -smooth costs, the step size is less than . However, many works have observed that in machine learning applications step sizes often do not fulfill this condition, yet (stochastic) gradient descent still converges, albeit in an unstable manner. We investigate this unstable convergence phenomenon from first principles, and discuss key causes behind it. We also identify its main characteristics, and how they interrelate based on both theory and experiments, offering a principled view toward understanding the phenomenon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of StabilityZixuan Wang, Zhouzi Li, Jian LiNeurIPS 2022 · 被引用 71 次
- Implicit Bias of the Step Size in Linear Diagonal Neural NetworksMor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, Daniel SoudryICML 2022 · 被引用 57 次
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 被引用 46 次
它引用的顶会 Paper1
相关 Paper
- On the Unstable Convergence Regime of Gradient DescentShuo Chen, Jiaying Peng, Xiaolong Li, Yao ZhaoAAAI 2024 · 被引用 1 次
- Beyond the Edge of Stability via Two-step Gradient UpdatesLei Chen, Joan BrunaICML 2023 · 被引用 22 次
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt 等ICML 2026
- Special Properties of Gradient Descent with Large Learning RatesAmirkeivan Mohtashami, Martin Jaggi, Sebastian U. StichICML 2023 · 被引用 16 次
- Constant Stepsize Local GD for Logistic Regression: Acceleration by InstabilityMichael Crawshaw, Blake Woodworth, Mingrui LiuICML 2025
