Embedding Principle of Loss Landscape of Deep Neural Networks
Yaoyu Zhang, Zhongwang Zhang, Tao Luo, Zhi-Qin John Xu
摘要
Understanding the structure of loss landscape of deep neural networks (DNNs) is obviously important. In this work, we prove an embedding principle that the loss landscape of a DNN "contains" all the critical points of all the narrower DNNs. More precisely, we propose a critical embedding such that any critical point, e.g., local or global minima, of a narrower DNN can be embedded to a critical point/affine subspace of the target DNN with higher degeneracy and preserving the DNN output function. Note that, given any training data, differentiable loss function and differentiable activation function, this embedding structure of critical points holds. This general structure of DNNs is starkly different from other nonconvex problems such as protein-folding. Empirically, we find that a wide DNN is often attracted by highly-degenerate critical points that are embedded from narrow DNNs. The embedding principle provides a new perspective to study the general easy optimization of wide DNNs and unravels a potential implicit low-complexity regularization during the training. Overall, our work provides a skeleton for the study of loss landscape of DNNs and its implication, by which a more exact and comprehensive understanding can be anticipated in the near future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Towards Understanding the Condensation of Neural Networks at Initial TrainingHanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang 等NeurIPS 2022 · 被引用 42 次
- Empirical Phase Diagram for Three-layer Neural Networks with Infinite WidthHanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo 等NeurIPS 2022 · 被引用 22 次
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 被引用 15 次
- Flat Channels to Infinity in Neural Loss LandscapesFlavio Martinelli, Alexander van Meegen, Berfin Simsek, Wulfram Gerstner 等NeurIPS 2025 · 被引用 6 次
- Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix CompletionZhiwei Bai, Jiajie Zhao, Yaoyu ZhangNeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper3
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro 等ICML 2021 · 被引用 136 次
- A Dynamical Central Limit Theorem for Shallow Neural NetworksZhengdao Chen, Grant M. Rotskoff, Joan Bruna, Eric Vanden-EijndenNeurIPS 2020 · 被引用 33 次
相关 Paper
- Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network EmbeddingZhengqing Wu, Berfin Simsek, François Gaston GedICLR 2025
- A Deeper Look at the Hessian Eigenspectrum of Deep Neural Networks and its Applications to RegularizationAdepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, Vineeth N. BalasubramanianAAAI 2021 · 被引用 60 次
- The Persistence of Neural Collapse Despite Low-Rank BiasConnall Garrod, Jonathan P. KeatingNeurIPS 2025 · 被引用 2 次
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 被引用 33 次
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
