Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
Zhengqing Wu, Berfin Simsek, François Gaston Ged
Abstract
In this paper, we investigate the loss landscape of one-hidden-layer neural networks with ReLUlike activation functions trained with the empirical squared loss. As the activation function is non-differentiable, it is so far unclear how to completely characterize the stationary points. We propose the conditions for stationarity that apply to both non-differentiable and differentiable cases. Additionally, we show that, if a stationary point does not contain "escape neurons", which are defined with first-order conditions, then it must be a local minimum. Moreover, for the scalar-output case, the presence of an escape neuron guarantees that the stationary point is not a local minimum. Our results refine the description of the saddle-to-saddle training process starting from infinitesimally small (vanishing) initialization for shallow ReLU-like networks, linking saddle escaping directly with the parameter changes of escape neurons. Moreover, we are also able to fully discuss how network embedding, which is to instantiate a narrower network within a wider network, reshapes the stationary points.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06d5f273-a9aa-49cb-81ee-dadb4bede125Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
- Saddle-to-Saddle Dynamics in Diagonal Linear NetworksScott Pesme, Nicolas FlammarionNeurIPS 2023 · 68 citations
- Embedding Principle of Loss Landscape of Deep Neural NetworksYaoyu Zhang, Zhongwang Zhang, Tao Luo, Zhi-Qin John XuNeurIPS 2021 · 48 citations
- Learning a Neuron by a Shallow ReLU Network: Dynamics and Implicit Bias for Correlated InputsDmitry Chistikov, Matthias Englert, Ranko LazicNeurIPS 2023 · 22 citations
Related papers
- Topological obstruction to the training of shallow ReLU neural networksMarco Nurisso, Pierrick Leroy, Francesco VaccarinoNeurIPS 2024 · 6 citations
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 33 citations
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 1 citation
- The Implicit Bias of Minima Stability in Multivariate Shallow ReLU NetworksMor Shpigel Nacson, Rotem Mulayoff, Greg Ongie, Tomer Michaeli et al.ICLR 2023 · 3 citations
- Linear Regularizers Enforce the Strict Saddle PropertyMatthew Ubl, Matthew Hale, Kasra YazdaniAAAI 2023 · 3 citations
