Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka Zdeborová
摘要
We study the dynamics of optimization and the generalization properties of one-hidden layer neural networks with quadratic activation function in the over-parametrized regime where the layer width is larger than the input dimension . We consider a teacher-student scenario where the teacher has the same structure as the student with a hidden layer of smaller width . We describe how the empirical loss landscape is affected by the number of data samples and the width of the teacher network. In particular we determine how the probability that there be no spurious minima on the empirical loss depends on , , and , thereby establishing conditions under which the neural network can in principle recover the teacher. We also show that under the same conditions gradient descent dynamics on the empirical loss converges and leads to small generalization error, i.e. it enables recovery in practice. Finally we characterize the time-convergence rate of gradient descent in the limit of a large number of samples. These results are confirmed by numerical experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 被引用 86 次
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 被引用 43 次
- On the Cryptographic Hardness of Learning Single Periodic NeuronsMin Jae Song, Ilias Zadik, Joan BrunaNeurIPS 2021 · 被引用 39 次
- A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNsGadi Naveh, Zohar RingelNeurIPS 2021 · 被引用 38 次
- Emergence and scaling laws in SGD learning of shallow neural networksYunwei Ren, Eshaan Nichani, Denny Wu, Jason D. LeeNeurIPS 2025 · 被引用 33 次
它引用的顶会 Paper2
- Bad Global Minima Exist and SGD Can Reach ThemShengchao Liu, Dimitris S. Papailiopoulos, Dimitris AchlioptasNeurIPS 2020 · 被引用 89 次
- Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase RetrievalStefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala 等NeurIPS 2020 · 被引用 32 次
相关 Paper
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 被引用 53 次
- On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student SettingShunta Akiyama, Taiji SuzukiICML 2021 · 被引用 16 次
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 被引用 1 次
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala 等NeurIPS 2022 · 被引用 59 次
- Rethinking Gauss-Newton for learning over-parameterized modelsMichael Arbel, Romain Menegaux, Pierre WolinskiNeurIPS 2023 · 被引用 10 次
