How Gradient descent balances features: A dynamical analysis for two-layer neural networks
Zhenyu Zhu, Fanghui Liu, Volkan Cevher
Abstract
This paper investigates the fundamental regression task of learning k neurons (a.k.a. teachers) from Gaussian input, using two-layer ReLU neural networks with width m (a.k.a. students) and m, k = O(1), trained via gradient descent under proper initialization and a small step-size. Our analysis follows a threephase structure: alignment after weak recovery, tangential growth, and local convergence, providing deeper insights into the learning dynamics of gradient descent (GD). We prove the global convergence at the rate of O(T -3 ) for the zero loss of excess risk. Additionally, our results show that GD automatically groups and balances student neurons, revealing an implicit bias toward achieving the minimum "balanced" ℓ 2 -norm in the solution. Our work extends beyond previous studies in exact-parameterization setting (m = k = 1, (Yehudai and Ohad, 2020)) and single-neuron setting (m ≥ k = 1, (Xu and Du, 2023)). The key technical challenge lies in handling the interactions between multiple teachers and students during training, which we address by refining the alignment analysis in Phase 1 and introducing a new dynamic system analysis for tangential components in Phase 2. Our results pave the way for further research on optimizing neural network training dynamics and understanding implicit biases in more complex architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcb347c3-2d5c-43c4-9f5d-bbe342aa9405Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 139 citations
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
- The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap ExponentsYatin Dandi, Emanuele Troiani, Luca Arnaboldi, Luca Pesce et al.ICML 2024 · 41 citations
- Early Neuron Alignment in Two-layer ReLU Networks with Small InitializationHancheng Min, Enrique Mallada, René VidalICLR 2024 · 31 citations
Related papers
- Learning a Neuron by a Shallow ReLU Network: Dynamics and Implicit Bias for Correlated InputsDmitry Chistikov, Matthias Englert, Ranko LazicNeurIPS 2023 · 22 citations
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 26 citations
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 17 citations
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 1 citation
- Neural Networks Efficiently Learn Low-Dimensional Representations with SGDAlireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas et al.ICLR 2023 · 6 citations
