Stepping on the Edge: Curvature Aware Learning Rate Tuners
Vincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz, Mathieu Blondel, Fabian Pedregosa
Abstract
Curvature information -- particularly, the largest eigenvalue of the loss Hessian, known as the sharpness -- often forms the basis for learning rate tuners. However, recent work has shown that the curvature information undergoes complex dynamics during training, going from a phase of increasing sharpness to eventual stabilization. We analyze the closed-loop feedback effect between learning rate tuning and curvature. We find that classical learning rate tuners may yield greater one-step loss reduction, yet they ultimately underperform in the long term when compared to constant learning rates in the full batch regime. These models break the stabilization of the sharpness, which we explain using a simplified model of the joint dynamics of the learning rate and the curvature. To further investigate these effects, we introduce a new learning rate tuning method, Curvature Dynamics Aware Tuning (CDAT), which prioritizes long term curvature stabilization over instantaneous progress on the objective. In the full batch regime, CDAT shows behavior akin to prefixed warm-up schedules on deep learning objectives, outperforming tuned constant learning rates. In the mini batch regime, we observe that stochasticity introduces confounding effects that explain the previous success of some learning rate tuners at appropriate batch sizes. Our findings highlight the critical role of understanding the joint dynamics of the learning rate and curvature, beyond greedy minimization, to diagnose failures and design effective adaptive learning rate tuners.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Why Do We Need Warm-up? A Theoretical PerspectiveFoivos Alimisis, Rustem Islamov, Aurelien LucchiICML 2026 · 8 citations
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network TrainingWenjie Zhou, Bohan Wang, Wei Chen, Xueqi ChengEMNLP 2025
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt et al.ICML 2026
- Exploiting weight-space symmetries for approximating curvatureArtem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu et al.ICML 2026
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang et al.ICLR 2024 · 264 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
Related papers
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 21 citations
- A Loss Curvature Perspective on Training Instabilities of Deep Learning ModelsJustin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta et al.ICLR 2022 · 49 citations
- Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian SpectrumsRui Pan, Haishan Ye, Tong ZhangICLR 2022 · 20 citations
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 87 citations
- Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of StabilityAlex Damian, Eshaan Nichani, Jason D. LeeICLR 2023 · 3 citations
