Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods
Shunta Akiyama, Taiji Suzuki
摘要
While deep learning has outperformed other methods for various tasks, theoretical frameworks that explain its reason have not been fully established. To address this issue, we investigate the excess risk of two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown teacher network through its outputs. Especially, we consider the student network that has the same width as the teacher network and is trained in two phases: first by noisy gradient descent and then by the vanilla gradient descent. Our result shows that the student network provably reaches a near-global optimal solution and outperforms any kernel methods estimator (more generally, linear estimators), including neural tangent kernel approach, random feature model, and other kernel methods, in a sense of the minimax optimal rate. The key concept inducing this superiority is the non-convexity of the neural network models. Even though the loss landscape is highly non-convex, the student network adaptively learns the teacher neurons.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Provable Multi-Task Representation Learning by Two-Layer ReLU Neural NetworksLiam Collins, Hamed Hassani, Mahdi Soltanolkotabi, Aryan Mokhtari 等ICML 2024 · 被引用 15 次
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 被引用 15 次
- Uniform-in-Time Wasserstein Stability Bounds for (Noisy) Stochastic Gradient DescentLingjiong Zhu, Mert Gürbüzbalaban, Anant Raj, Umut SimsekliNeurIPS 2023 · 被引用 10 次
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos 等ICLR 2024 · 被引用 2 次
- How Gradient descent balances features: A dynamical analysis for two-layer neural networksZhenyu Zhu, Fanghui Liu, Volkan CevherICLR 2025
相关 Paper
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methodsTaiji Suzuki, Shunta AkiyamaICLR 2021 · 被引用 12 次
- On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student SettingShunta Akiyama, Taiji SuzukiICML 2021 · 被引用 16 次
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 被引用 15 次
- Optimization and Adaptive Generalization of Three layer Neural NetworksKhashayar Gatmiry, Stefanie Jegelka, Jonathan A. KelnerICLR 2022 · 被引用 1 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
