Optimization and Generalization of Shallow Neural Networks with Quadratic Activation Functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka Zdeborová
Abstract
We study the dynamics of optimization and the generalization properties of one-hidden layer neural networks with quadratic activation function in the over-parametrized regime where the layer width is larger than the input dimension . We consider a teacher-student scenario where the teacher has the same structure as the student with a hidden layer of smaller width . We describe how the empirical loss landscape is affected by the number of data samples and the width of the teacher network. In particular we determine how the probability that there be no spurious minima on the empirical loss depends on , , and , thereby establishing conditions under which the neural network can in principle recover the teacher. We also show that under the same conditions gradient descent dynamics on the empirical loss converges and leads to small generalization error, i.e. it enables recovery in practice. Finally we characterize the time-convergence rate of gradient descent in the limit of a large number of samples. These results are confirmed by numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab24e65d-8eb0-4305-bedd-52ce296344c5Cited by top-tier papers25
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 43 citations
- On the Cryptographic Hardness of Learning Single Periodic NeuronsMin Jae Song, Ilias Zadik, Joan BrunaNeurIPS 2021 · 39 citations
- A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNsGadi Naveh, Zohar RingelNeurIPS 2021 · 38 citations
- Emergence and scaling laws in SGD learning of shallow neural networksYunwei Ren, Eshaan Nichani, Denny Wu, Jason D. LeeNeurIPS 2025 · 33 citations
Builds on2
- Bad Global Minima Exist and SGD Can Reach ThemShengchao Liu, Dimitris S. Papailiopoulos, Dimitris AchlioptasNeurIPS 2020 · 89 citations
- Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase RetrievalStefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala et al.NeurIPS 2020 · 32 citations
Related papers
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
- On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student SettingShunta Akiyama, Taiji SuzukiICML 2021 · 16 citations
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 1 citation
- Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networksRodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala et al.NeurIPS 2022 · 59 citations
- Rethinking Gauss-Newton for learning over-parameterized modelsMichael Arbel, Romain Menegaux, Pierre WolinskiNeurIPS 2023 · 10 citations
