Annihilation of Spurious Minima in Two-Layer ReLU Networks
Yossi Arjevani, Michael Field
摘要
We study the optimization problem associated with fitting two-layer ReLU neural networks with respect to the squared loss, where labels are generated by a target network. Use is made of the rich symmetry structure to develop a novel set of tools for studying the mechanism by which over-parameterization annihilates spurious minima. Sharp analytic estimates are obtained for the loss and the Hessian spectrum at different minima, and it is proved that adding neurons can turn symmetric spurious minima into saddles; minima of lesser symmetry require more neurons. Using Cauchy's interlacing theorem, we prove the existence of descent directions in certain subspaces arising from the symmetry structure of the loss function. This analytic approach uses techniques, new to the field, from algebraic geometry, representation theory and symmetry breaking, and confirms rigorously the effectiveness of over-parameterization in making the associated loss landscape accessible to gradient-based methods. For a fixed number of neurons and inputs, the spectral results remain true under symmetry breaking perturbation of the target. This example highlights the special role that the standard representation plays in the annihilation of spurious minima (see Section 5 and the concluding remarks). The sharp estimates of the Hessian spectrum further demonstrate how symmetry breaking enables a complete characterization of the dynamics of gradient-based methods, locally, in the vicinity of symmetric critical points. The dependence of such methods on stability of critical points therefore indicates that attempts for a global theory should be preceded by a good description of the mechanism by which spurious minima transform into saddles-the aim of this work. Next, we relate our results to the existing literature. Annihilation of spurious minima on account of over-parameterization. Existing methods for the analysis of optimization problem (2) include: mean-field [4], optimal transport [2], NTK [20, 21, 22] and the thermodynamic limit [5, 16, 23, 24] . These methods operate by passing to limiting regimes where the number of inputs or neurons is taken to infinity. A growing number of works has limited the explanatory power of such approaches [25, 26] . Approaches for addressing the loss landscapes in finite parameter regimes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 被引用 217 次
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 被引用 22 次
- On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student SettingShunta Akiyama, Taiji SuzukiICML 2021 · 被引用 16 次
- Student Specialization in Deep Rectified Networks With Finite Width and Input DimensionYuandong TianICML 2020 · 被引用 12 次
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry IIYossi Arjevani, Michael FieldNeurIPS 2021
相关 Paper
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 被引用 1 次
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro 等ICML 2021 · 被引用 136 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Simplicity Bias and Optimization Threshold in Two-Layer ReLU NetworksEtienne Boursier, Nicolas FlammarionICML 2025
