Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, Lenka Zdeborová
Abstract
Despite the widespread use of gradient-based algorithms for optimizing high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retrieval from random measurements. When the ratio of the number of measurements over the input dimension is small the dynamics remains trapped in spurious minima with large basins of attraction. We find analytically that above a critical ratio those critical points become unstable developing a negative direction toward the signal. By numerical experiments we show that in this regime the gradient flow algorithm is not trapped; it drifts away from the spurious critical points along the unstable direction and succeeds in finding the global minimum. Using tools from statistical physics we characterize this phenomenon, which is related to a BBP-type transition in the Hessian of the spurious minima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a3e8d69-a9af-4e25-a819-ed505fe20165Cited by top-tier papers9
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 65 citations
- An Analytical Theory of Curriculum Learning in Teacher-Student NetworksLuca Saglietti, Stefano Sarao Mannelli, Andrew M. SaxeNeurIPS 2022 · 43 citations
- On the Cryptographic Hardness of Learning Single Periodic NeuronsMin Jae Song, Ilias Zadik, Joan BrunaNeurIPS 2021 · 39 citations
- Bayes-optimal learning of an extensive-width neural network from quadratically many samplesAntoine Maillard, Emanuele Troiani, Simon Martin, Florent Krzakala et al.NeurIPS 2024 · 26 citations
- Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex ProblemsStefano Sarao Mannelli, Pierfrancesco UrbaniNeurIPS 2021 · 12 citations
Builds on1
Related papers
- Overparametrization bends the landscape: BBP transitions at initialization in simple Neural NetworksBrandon Livio Annesi, Dario Bocchi, Chiara CammarotaICLR 2026 · 3 citations
- A Continuous-Time Mirror Descent Approach to Sparse Phase RetrievalFan Wu, Patrick RebeschiniNeurIPS 2020 · 16 citations
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 96 citations
- Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstructionDominik Stöger, Mahdi SoltanolkotabiNeurIPS 2021 · 101 citations
- From Gradient Flow on Population Loss to Learning with Stochastic Gradient DescentChristopher De Sa, Satyen Kale, Jason D. Lee, Ayush Sekhari et al.NeurIPS 2022 · 5 citations
