Understanding the Dynamics of Gradient Flow in Overparameterized Linear models
Salma Tarmoun, Guilherme França, Benjamin D. Haeffele, René Vidal
摘要
We provide a detailed analysis of the dynamics of the gradient flow in overparameterized two-layer linear models. A particularly interesting feature of this model is that its nonlinear dynamics can be exactly solved as a consequence of a large number of conservation laws that constrain the system to follow particular trajectories. More precisely, the gradient flow preserves the difference of the Gramian matrices of the input and output weights, and its convergence to equilibrium depends on both the magnitude of that difference (which is fixed at initialization) and the spectrum of the data. In addition, and generalizing prior work, we prove our results without assuming small, balanced or spectral initialization for the weights. Moreover, we establish interesting mathematical connections between matrix factorization problems and differential equations of the Riccati type.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Exact learning dynamics of deep linear networks with prior knowledgeLukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. SaxeNeurIPS 2022 · 被引用 75 次
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 被引用 54 次
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 等NeurIPS 2024 · 被引用 48 次
- Symmetry Teleportation for Accelerated OptimizationBo Zhao, Nima Dehmamy, Robin Walters, Rose YuNeurIPS 2022 · 被引用 33 次
- On the spectral bias of two-layer linear networksAditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2023 · 被引用 27 次
相关 Paper
- On the Convergence of Gradient Flow on Multi-layer Linear ModelsHancheng Min, René Vidal, Enrique MalladaICML 2023 · 被引用 13 次
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 被引用 47 次
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- Keep the Momentum: Conservation Laws beyond Euclidean Gradient FlowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2024 · 被引用 8 次
- Intrinsic training dynamics of deep neural networksSibylle Marcotte, Gabriel Peyré, Rémi GribonvalICLR 2026 · 被引用 4 次
