Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, Lenka Zdeborová
摘要
Despite the widespread use of gradient-based algorithms for optimizing high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retrieval from random measurements. When the ratio of the number of measurements over the input dimension is small the dynamics remains trapped in spurious minima with large basins of attraction. We find analytically that above a critical ratio those critical points become unstable developing a negative direction toward the signal. By numerical experiments we show that in this regime the gradient flow algorithm is not trapped; it drifts away from the spurious critical points along the unstable direction and succeeds in finding the global minimum. Using tools from statistical physics we characterize this phenomenon, which is related to a BBP-type transition in the Hessian of the spurious minima.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 被引用 65 次
- An Analytical Theory of Curriculum Learning in Teacher-Student NetworksLuca Saglietti, Stefano Sarao Mannelli, Andrew M. SaxeNeurIPS 2022 · 被引用 43 次
- On the Cryptographic Hardness of Learning Single Periodic NeuronsMin Jae Song, Ilias Zadik, Joan BrunaNeurIPS 2021 · 被引用 39 次
- Bayes-optimal learning of an extensive-width neural network from quadratically many samplesAntoine Maillard, Emanuele Troiani, Simon Martin, Florent Krzakala 等NeurIPS 2024 · 被引用 26 次
- Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex ProblemsStefano Sarao Mannelli, Pierfrancesco UrbaniNeurIPS 2021 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- Overparametrization bends the landscape: BBP transitions at initialization in simple Neural NetworksBrandon Livio Annesi, Dario Bocchi, Chiara CammarotaICLR 2026 · 被引用 3 次
- A Continuous-Time Mirror Descent Approach to Sparse Phase RetrievalFan Wu, Patrick RebeschiniNeurIPS 2020 · 被引用 16 次
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 被引用 96 次
- Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstructionDominik Stöger, Mahdi SoltanolkotabiNeurIPS 2021 · 被引用 101 次
- From Gradient Flow on Population Loss to Learning with Stochastic Gradient DescentChristopher De Sa, Satyen Kale, Jason D. Lee, Ayush Sekhari 等NeurIPS 2022 · 被引用 5 次
