SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
Julian Kranz, Davide Gallon, Steffen Dereich, Arnulf Jentzen
摘要
We study gradient flows for loss landscapes of fully connected feedforward neural networks with commonly used continuously differentiable activation functions such as the logistic, hyperbolic tangent, softplus or GELU function. We prove that the gradient flow either converges to a critical point or diverges to infinity while the loss converges to an asymptotic critical value. Moreover, we prove the existence of a threshold such that the loss value of any gradient flow initialized at most above the optimal level converges to it. For polynomial target functions and sufficiently big architecture and data set, we prove that the optimal loss value is zero and can only be realized asymptotically. From this setting, we deduce our main result that any gradient flow with sufficiently good initialization diverges to infinity. Our proof heavily relies on the geometry of o-minimal structures. We confirm these theoretical findings with numerical experiments and extend our investigation to more realistic scenarios, where we observe an analogous behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin 等NeurIPS 2023 · 被引用 93 次
相关 Paper
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Topological obstruction to the training of shallow ReLU neural networksMarco Nurisso, Pierrick Leroy, Francesco VaccarinoNeurIPS 2024 · 被引用 6 次
- Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic ActivationsAlexandru Craciun, Debarghya GhoshdastidarNeurIPS 2025 · 被引用 1 次
- Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network EmbeddingZhengqing Wu, Berfin Simsek, François Gaston GedICLR 2025
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 被引用 51 次
