Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks
Alexander Shevchenko, Marco Mondelli
摘要
The optimization of multilayer neural networks typically leads to a solution with zero training error, yet the landscape can exhibit spurious local minima and the minima can be disconnected. In this paper, we shed light on this phenomenon: we show that the combination of stochastic gradient descent (SGD) and over-parameterization makes the landscape of multilayer neural networks approximately connected and thus more favorable to optimization. More specifically, we prove that SGD solutions are connected via a piecewise linear path, and the increase in loss along this path vanishes as the number of neurons grows large. This result is a consequence of the fact that the parameters found by SGD are increasingly dropout stable as the network becomes wider. We show that, if we remove part of the neurons (and suitably rescale the remaining ones), the change in loss is independent of the total number of neurons, and it depends only on how many neurons are left. Our results exhibit a mild dependence on the input dimension: they are dimension-free for two-layer networks and require the number of neurons to scale linearly with the dimension for multilayer networks. We validate our theoretical findings with numerical experiments for different architectures and classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural NetworksAlexandra Peste, Eugenia Iofinova, Adrian Vladu, Dan AlistarhNeurIPS 2021 · 被引用 84 次
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou 等NeurIPS 2021 · 被引用 51 次
- Deep Networks on Toroids: Removing Symmetries Reveals the Structure of Flat Regions in the Landscape GeometryFabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer 等ICML 2022 · 被引用 30 次
- Global Convergence of Three-layer Neural Networks in the Mean Field RegimeHuy Tuan Pham, Phan-Minh NguyenICLR 2021 · 被引用 23 次
- Redundant representations help generalization in wide neural networksDiego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro LaioNeurIPS 2022 · 被引用 13 次
相关 Paper
- When Are Solutions Connected in Deep Networks?Quynh Nguyen, Pierre Bréchet, Marco MondelliNeurIPS 2021 · 被引用 12 次
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networksJie Huang, Bruno Loureiro, Stefano Sarao MannelliICML 2026 · 被引用 1 次
- Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural NetworksAmit Peleg, Matthias HeinICML 2024
- Exact Solutions of a Deep Linear NetworkLiu Ziyin, Botao Li, Xiangming MengNeurIPS 2022 · 被引用 31 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
