Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rank
Zihan Wang, Arthur Jacot
Abstract
The -regularized loss of Deep Linear Networks (DLNs) with more than one hidden layers has multiple local minima, corresponding to matrices with different ranks. In tasks such as matrix completion, the goal is to converge to the local minimum with the smallest rank that still fits the training data. While rank-underestimating minima can be avoided since they do not fit the data, GD might get stuck at rank-overestimating minima. We show that with SGD, there is always a probability to jump from a higher rank minimum to a lower rank one, but the probability of jumping back is zero. More precisely, we define a sequence of sets so that contains all minima of rank or less (and not more) that are absorbing for small enough ridge parameters and learning rates : SGD has prob. 0 of leaving , and from any starting point there is a non-zero prob. for SGD to go in .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa89b396-977d-40e6-93dc-e939a2ebcfa5Cited by top-tier papers18
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 52 citations
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
- Mixed Dynamics In Linear Networks: Unifying the Lazy and Active RegimesZhenfeng Tu, Santiago Aranguri, Arthur JacotNeurIPS 2024 · 18 citations
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 14 citations
- AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMsDi He, Songjun Tu, Ajay Jaiswal, Li Shen et al.NeurIPS 2025 · 14 citations
Builds on12
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
Related papers
- Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle EscapeIoannis Bantzis, James B. Simon, Arthur JacotICLR 2026 · 4 citations
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 44 citations
- Bad Global Minima Exist and SGD Can Reach ThemShengchao Liu, Dimitris S. Papailiopoulos, Dimitris AchlioptasNeurIPS 2020 · 89 citations
- The Global Convergence Time of Stochastic Gradient Descent in Non-Convex Landscapes: Sharp Estimates via Large DeviationsWaïss Azizian, Franck Iutzeler, Jérôme Malick, Panayotis MertikopoulosICML 2025
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 5 citations
