A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, Hossein Mobahi
摘要
We study the implicit bias of gradient flow (i.e., gradient descent with infinitesimal step size) on linear neural network training. We propose a tensor formulation of neural networks that includes fully-connected, diagonal, and convolutional networks as special cases, and investigate the linear version of the formulation called linear tensor networks. With this formulation, we can characterize the convergence direction of the network parameters as singular vectors of a tensor defined by the network. For L-layer linear tensor networks that are orthogonally decomposable, we show that gradient flow on separable classification finds a stationary point of the 2/L max-margin problem in a "transformed" input space defined by the network. For underdetermined regression, we prove that gradient flow finds a global minimum which minimizes a norm-like function that interpolates between weighted 1 and 2 norms in the transformed input space. Our theorems subsume existing results in the literature while removing standard convergence assumptions. We also provide experiments that corroborate our analysis. * Based on work performed during internship at Google Research
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper48
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 110 次
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
- Generalization on the Unseen, Logic Reasoning and Degree CurriculumEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Kevin RizkICML 2023 · 被引用 68 次
它引用的顶会 Paper6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear NetworksWei Hu, Lechao Xiao, Jeffrey PenningtonICLR 2020 · 被引用 136 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- Identity Crisis: Memorization and Generalization Under Extreme OverparameterizationChiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C. Mozer 等ICLR 2020 · 被引用 96 次
相关 Paper
- Implicit Regularization in Tensor FactorizationNoam Razin, Asaf Maman, Nadav CohenICML 2021 · 被引用 60 次
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro 等ICLR 2023 · 被引用 5 次
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 被引用 24 次
- Implicit Regularization for Tubal Tensor Factorizations via Gradient DescentSanthosh Karnik, Anna Veselovska, Mark A. Iwen, Felix KrahmerICML 2025
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 被引用 39 次
