On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting
Shunta Akiyama, Taiji Suzuki
Abstract
Deep learning empirically achieves high performance in many applications, but its training dynamics has not been fully understood theoretically. In this paper, we explore theoretical analysis on training two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown teacher network through its outputs. We show that with a specific regularization and sufficient over-parameterization, the student network can identify the parameters of the teacher network with high probability via gradient descent with a norm dependent stepsize even though the objective function is highly non-convex. The key theoretical tool is the measure representation of the neural networks and a novel application of a dual certificate argument for sparse estimation on a measure space. We analyze the global minima and global convergence property in the measure space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 15 citations
- Annihilation of Spurious Minima in Two-Layer ReLU NetworksYossi Arjevani, Michael FieldNeurIPS 2022 · 14 citations
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 14 citations
- Neural Networks Efficiently Learn Low-Dimensional Representations with SGDAlireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas et al.ICLR 2023 · 6 citations
- How does Gradient Descent Learn Features - A Local Analysis for Regularized Two-Layer Neural NetworksMo Zhou, Rong GeNeurIPS 2024 · 5 citations
Related papers
- Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel MethodsShunta Akiyama, Taiji SuzukiICLR 2023 · 1 citation
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 65 citations
- Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic ActivationsAlexandru Craciun, Debarghya GhoshdastidarNeurIPS 2025 · 1 citation
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 15 citations
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry et al.NeurIPS 2024 · 15 citations
