Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
Yossi Arjevani, Michael Field
Abstract
We consider the optimization problem associated with fitting two-layers ReLU networks with respect to the squared loss, where labels are generated by a target network. We leverage the rich symmetry structure to analytically characterize the Hessian at various families of spurious minima in the natural regime where the number of inputs and the number of hidden neurons is finite. In particular, we prove that for standard Gaussian inputs: (a) of the eigenvalues of the Hessian, concentrate near zero, (b) of the eigenvalues grow linearly with . Although this phenomenon of extremely skewed spectrum has been observed many times before, to our knowledge, this is the first time it has been established rigorously. Our analytic approach uses techniques, new to the field, from symmetry breaking and representation theory, and carries important implications for our ability to argue about statistical generalization through local curvature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e81a3c27-3eb6-4d16-8636-592f40479fecCited by top-tier papers8
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 26 citations
- Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian SpectrumsRui Pan, Haishan Ye, Tong ZhangICLR 2022 · 20 citations
- Annihilation of Spurious Minima in Two-Layer ReLU NetworksYossi Arjevani, Michael FieldNeurIPS 2022 · 14 citations
- A Classification of -invariant Shallow Neural NetworksDevanshu Agrawal, James OstrowskiNeurIPS 2022 · 11 citations
- Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient NoiseRui Pan, Yuxing Liu, Xiaoyu Wang, Tong ZhangICLR 2024 · 10 citations
Related papers
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry IIYossi Arjevani, Michael FieldNeurIPS 2021
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 1 citation
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 8 citations
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun et al.ICML 2025
