How Gradient descent balances features: A dynamical analysis for two-layer neural networks
Zhenyu Zhu, Fanghui Liu, Volkan Cevher
摘要
This paper investigates the fundamental regression task of learning k neurons (a.k.a. teachers) from Gaussian input, using two-layer ReLU neural networks with width m (a.k.a. students) and m, k = O(1), trained via gradient descent under proper initialization and a small step-size. Our analysis follows a threephase structure: alignment after weak recovery, tangential growth, and local convergence, providing deeper insights into the learning dynamics of gradient descent (GD). We prove the global convergence at the rate of O(T -3 ) for the zero loss of excess risk. Additionally, our results show that GD automatically groups and balances student neurons, revealing an implicit bias toward achieving the minimum "balanced" ℓ 2 -norm in the solution. Our work extends beyond previous studies in exact-parameterization setting (m = k = 1, (Yehudai and Ohad, 2020)) and single-neuron setting (m ≥ k = 1, (Xu and Du, 2023)). The key technical challenge lies in handling the interactions between multiple teachers and students during training, which we address by refining the alignment analysis in Phase 1 and introducing a new dynamic system analysis for tangential components in Phase 2. Our results pave the way for further research on optimizing neural network training dynamics and understanding implicit biases in more complex architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap ExponentsYatin Dandi, Emanuele Troiani, Luca Arnaboldi, Luca Pesce 等ICML 2024 · 被引用 41 次
- Early Neuron Alignment in Two-layer ReLU Networks with Small InitializationHancheng Min, Enrique Mallada, René VidalICLR 2024 · 被引用 31 次
相关 Paper
- Learning a Neuron by a Shallow ReLU Network: Dynamics and Implicit Bias for Correlated InputsDmitry Chistikov, Matthias Englert, Ranko LazicNeurIPS 2023 · 被引用 22 次
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 被引用 26 次
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 被引用 17 次
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 被引用 1 次
- Neural Networks Efficiently Learn Low-Dimensional Representations with SGDAlireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas 等ICLR 2023 · 被引用 6 次
