Deep linear networks for regression are implicitly regularized towards flat minima
Pierre Marion, Lénaïc Chizat
摘要
The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small one. Indeed, we show a lower bound on the sharpness of minimizers, which grows linearly with depth. We then study the properties of the minimizer found by gradient flow, which is the limit of gradient descent with vanishing learning rate. We show an implicit regularization towards flat minima: the sharpness of the minimizer is no more than a constant times the lower bound. The constant depends on the condition number of the data covariance matrix, but not on width or depth. This result is proven both for a small-scale initialization and a residual initialization. Results of independent interest are shown in both cases. For small-scale initialization, we show that the learned weight matrices are approximately rank-one and that their singular vectors align. For residual initialization, convergence of the gradient flow for a Gaussian initialization of the residual network is proven. Numerical experiments illustrate our results and connect them to gradient descent with non-vanishing learning rate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networksLénaïc Chizat, Praneeth NetrapalliNeurIPS 2024 · 被引用 12 次
- FACT: a first-principles alternative to the Neural Feature Ansatz for how networks learn representationsEnric Boix Adserà, Neil Mallinar, James B. Simon, Misha BelkinICLR 2026 · 被引用 2 次
- Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?Tom Jacobs, Chao Zhou, Rebekka BurkholzICML 2025
- The Optimization Landscape of SGD Across the Feature Learning StrengthAlexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz PehlevanICLR 2025
- Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-PathwayHee-Sung Kim, Sungyoon LeeICML 2026
它引用的顶会 Paper20
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 被引用 173 次
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
相关 Paper
- Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to ChaosDayal Singh Kalra, Tianyu He, Maissam BarkeshliICLR 2025
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou 等ICLR 2023 · 被引用 1 次
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 被引用 43 次
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 被引用 21 次
