Understanding the Dynamics of Gradient Flow in Overparameterized Linear models
Salma Tarmoun, Guilherme França, Benjamin D. Haeffele, René Vidal
Abstract
We provide a detailed analysis of the dynamics of the gradient flow in overparameterized two-layer linear models. A particularly interesting feature of this model is that its nonlinear dynamics can be exactly solved as a consequence of a large number of conservation laws that constrain the system to follow particular trajectories. More precisely, the gradient flow preserves the difference of the Gramian matrices of the input and output weights, and its convergence to equilibrium depends on both the magnitude of that difference (which is fixed at initialization) and the spectrum of the data. In addition, and generalizing prior work, we prove our results without assuming small, balanced or spectral initialization for the weights. Moreover, we establish interesting mathematical connections between matrix factorization problems and differential equations of the Riccati type.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9b9cc55-d326-483e-a6c2-355960b46f3cCited by top-tier papers23
- Exact learning dynamics of deep linear networks with prior knowledgeLukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. SaxeNeurIPS 2022 · 75 citations
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 54 citations
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen et al.NeurIPS 2024 · 48 citations
- Symmetry Teleportation for Accelerated OptimizationBo Zhao, Nima Dehmamy, Robin Walters, Rose YuNeurIPS 2022 · 33 citations
- On the spectral bias of two-layer linear networksAditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2023 · 27 citations
Related papers
- On the Convergence of Gradient Flow on Multi-layer Linear ModelsHancheng Min, René Vidal, Enrique MalladaICML 2023 · 13 citations
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 47 citations
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
- Keep the Momentum: Conservation Laws beyond Euclidean Gradient FlowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2024 · 8 citations
- Intrinsic training dynamics of deep neural networksSibylle Marcotte, Gabriel Peyré, Rémi GribonvalICLR 2026 · 4 citations
