Flatland: The Adventures of Gradient Descent with Large Step Sizes
Leonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt, Holger Rauhut
Abstract
The training of neural networks often entails objective functions that are not globally -smooth. For these functions, it is both theoretically and practically difficult to reply to the question: what is the largest possible step size that ensures the convergence of gradient descent (GD)? We address this longstanding open question in deep learning by providing a unifying definition of "large'' step sizes that requires only local Lipschitz (or even Hölder) continuity of the gradient. We design first-order adaptive methods that provably yield large step sizes and show that they operate at the edge of stability (EoS) right from the start of the training. In particular, the loss decreases nonmonotonically and the product between the step size and sharpness, i.e., the largest eigenvalue of the hessian, stays above the EoS threshold of 2 throughout training. Using our method, we are also able to minimize the sharpness all the way down to its global minimum. Contrary to expectation, we find that encountering globally-flat regions too early in the training may both slow down convergence and jeopardize the generalization ability of the network. Exploiting a self-stabilization argument, we allow GD to enter slightly sharper valleys and turn unsuccessful training runs into very successful ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62f6f159-b451-4a08-88e4-2d6fb63d4acdBuilds on16
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- Adaptive Gradient Descent without DescentYura Malitsky, Konstantin MishchenkoICML 2020 · 171 citations
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin et al.NeurIPS 2023 · 93 citations
- A Modern Look at the Relationship between Sharpness and GeneralizationMaksym Andriushchenko, Francesco Croce, Maximilian Müller, Matthias Hein et al.ICML 2023 · 92 citations
Related papers
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou et al.ICLR 2023 · 1 citation
- Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and BeyondItai Kreisler, Mor Shpigel Nacson, Daniel Soudry, Yair CarmonICML 2023 · 19 citations
- Non-Euclidean Gradient Descent Operates at the Edge of StabilityRustem Islamov, Michael Crawshaw, Jeremy Cohen, Robert GowerICML 2026 · 5 citations
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter et al.ICLR 2021 · 22 citations
- Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of StabilityAlex Damian, Eshaan Nichani, Jason D. LeeICLR 2023 · 3 citations
