Deep linear networks for regression are implicitly regularized towards flat minima
Pierre Marion, Lénaïc Chizat
Abstract
The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small one. Indeed, we show a lower bound on the sharpness of minimizers, which grows linearly with depth. We then study the properties of the minimizer found by gradient flow, which is the limit of gradient descent with vanishing learning rate. We show an implicit regularization towards flat minima: the sharpness of the minimizer is no more than a constant times the lower bound. The constant depends on the condition number of the data covariance matrix, but not on width or depth. This result is proven both for a small-scale initialization and a residual initialization. Results of independent interest are shown in both cases. For small-scale initialization, we show that the learned weight matrices are approximately rank-one and that their singular vectors align. For residual initialization, convergence of the gradient flow for a Gaussian initialization of the residual network is proven. Numerical experiments illustrate our results and connect them to gradient descent with non-vanishing learning rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db8507c2-4b2b-4c3e-95e0-4ac5f369af61Cited by top-tier papers5
- The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networksLénaïc Chizat, Praneeth NetrapalliNeurIPS 2024 · 12 citations
- FACT: a first-principles alternative to the Neural Feature Ansatz for how networks learn representationsEnric Boix Adserà, Neil Mallinar, James B. Simon, Misha BelkinICLR 2026 · 2 citations
- Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?Tom Jacobs, Chao Zhou, Rebekka BurkholzICML 2025
- The Optimization Landscape of SGD Across the Feature Learning StrengthAlexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz PehlevanICLR 2025
- Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-PathwayHee-Sung Kim, Sungyoon LeeICML 2026
Builds on20
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 155 citations
Related papers
- Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to ChaosDayal Singh Kalra, Tianyu He, Maissam BarkeshliICLR 2025
- Understanding Edge-of-Stability Training Dynamics with a Minimalist ExampleXingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou et al.ICLR 2023 · 1 citation
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 43 citations
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
- Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and widthDayal Singh Kalra, Maissam BarkeshliNeurIPS 2023 · 21 citations
