Piecewise linear activations substantially shape the loss surfaces of neural networks
Fengxiang He, Bohan Wang, Dacheng Tao
摘要
Understanding the loss surface of a neural network is fundamentally important to the understanding of deep learning. This paper presents how piecewise linear activation functions substantially shape the loss surfaces of neural networks. We first prove that the loss surfaces of many neural networks have infinite spurious local minima, which are defined as the local minima with higher empirical risks than the global minima. Our result holds for any neural network with arbitrary depth and arbitrary piecewise linear activation functions (excluding linear functions) under most loss functions in practice with some mild assumptions. This result demonstrates that the networks with piecewise linear activations possess substantial differences to the well-studied linear neural networks. Essentially, the underlying assumptions for the above result are consistent with most practical circumstances where the output layer is narrower than any hidden layer. In addition, the loss surface of a neural network with piecewise linear activations is partitioned into multiple smooth and multilinear cells by nondifferentiable boundaries. The constructed spurious local minima are concentrated in one cell as a valley: they are connected with each other by a continuous path, on which empirical risk is invariant. Further for one-hidden-layer networks, we prove that all local minima in a cell constitute an equivalence class; they are concentrated in a valley; and they are all global minima in the cell.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Modeling Image Composition for Complex Scene GenerationZuopeng Yang, Daqing Liu, Chaoyue Wang, Jie Yang 等CVPR 2022 · 被引用 37 次
- Truth or backpropaganda? An empirical investigation of deep learning theoryMicah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller 等ICLR 2020 · 被引用 36 次
- Exact Solutions of a Deep Linear NetworkLiu Ziyin, Botao Li, Xiangming MengNeurIPS 2022 · 被引用 31 次
- Auto Learning AttentionBenteng Ma, Jing Zhang, Yong Xia, Dacheng TaoNeurIPS 2020 · 被引用 28 次
- When Are Solutions Connected in Deep Networks?Quynh Nguyen, Pierre Bréchet, Marco MondelliNeurIPS 2021 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Pure and Spurious Critical Points: a Geometric Study of Linear NetworksMatthew Trager, Kathlén Kohn, Joan BrunaICLR 2020 · 被引用 41 次
- Empirical Studies on the Properties of Linear Regions in Deep Neural NetworksXiao Zhang, Dongrui WuICLR 2020 · 被引用 44 次
- Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network EmbeddingZhengqing Wu, Berfin Simsek, François Gaston GedICLR 2025
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
