Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width
Dayal Singh Kalra, Maissam Barkeshli
摘要
We systematically analyze optimization dynamics in deep neural networks (DNNs) trained with stochastic gradient descent (SGD) and study the effect of learning rate , depth , and width of the neural network. By analyzing the maximum eigenvalue of the Hessian of the loss, which is a measure of sharpness of the loss landscape, we find that the dynamics can show four distinct regimes: (i) an early time transient regime, (ii) an intermediate saturation regime, (iii) a progressive sharpening regime, and (iv) a late time edge of stability"regime. The early and intermediate regimes (i) and (ii) exhibit a rich phase diagram depending on $\eta \equiv c / \lambda_0^H $, $d$, and $w$. We identify several critical values of $c$, which separate qualitatively distinct phenomena in the early time dynamics of training loss and sharpness. Notably, we discover the opening up of a sharpness reduction"phase, where sharpness decreases at early times, as and are increased.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 被引用 87 次
- Catapults in SGD: spikes in the training loss and their impact on generalization through feature learningLibin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail BelkinICML 2024 · 被引用 29 次
- Super Consistency of Neural Network Landscapes and Learning Rate TransferLorenzo Noci, Alexandru Meterez, Thomas Hofmann, Antonio OrvietoNeurIPS 2024 · 被引用 25 次
- Non-Euclidean Gradient Descent Operates at the Edge of StabilityRustem Islamov, Michael Crawshaw, Jeremy Cohen, Robert GowerICML 2026 · 被引用 5 次
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
它引用的顶会 Paper21
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
相关 Paper
- Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of StabilityZixuan Wang, Zhouzi Li, Jian LiNeurIPS 2022 · 被引用 71 次
- Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to ChaosDayal Singh Kalra, Tianyu He, Maissam BarkeshliICLR 2025
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou 等ICLR 2023 · 被引用 1 次
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter 等ICLR 2021 · 被引用 22 次
- Deep linear networks for regression are implicitly regularized towards flat minimaPierre Marion, Lénaïc ChizatNeurIPS 2024 · 被引用 21 次
