Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data
Guillaume Braun, Bruno Loureiro, Minh Ha Quang, Masaaki Imaizumi
Abstract
Scaling laws describe how learning performance improves with data, compute, or training time, and have become a central theme in modern deep learning. We study this phenomenon in a canonical nonlinear model: phase retrieval with anisotropic Gaussian inputs whose covariance spectrum follows a power law. Unlike the isotropic case, where dynamics collapse to a two-dimensional system, anisotropy yields a qualitatively new regime in which an infinite hierarchy of coupled equations governs the evolution of the summary statistics. We develop a tractable reduction that reveals a three-phase trajectory: (i) fast escape from low alignment, (ii) slow convergence of the summary statistics, and (iii) spectral-tail learning in low-variance directions. From this decomposition, we derive explicit scaling laws for the mean-squared error, showing how spectral decay dictates convergence times and error curves. Experiments confirm the predicted phases and exponents. These results provide the first rigorous characterization of scaling laws in nonlinear regression with anisotropic data, highlighting how anisotropy reshapes learning dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c5af070-031e-48ce-8cda-0ece2e09d938Cited by top-tier papers2
- Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index LearningFilip Kovačević, Hong Chang Ji, Denny Wu, Mahdi Soltanolkotabi et al.ICML 2026 · 2 citations
- Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase RetrievalGuillaume Braun, Han Bao, Wei Huang, Masaaki ImaizumiICML 2026
Builds on14
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 109 citations
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 94 citations
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 84 citations
- 4+3 Phases of Compute-Optimal Neural Scaling LawsElliot Paquette, Courtney Paquette, Lechao Xiao, Jeffrey PenningtonNeurIPS 2024 · 70 citations
- Optimization and Generalization of Shallow Neural Networks with Quadratic Activation FunctionsStefano Sarao Mannelli, Eric Vanden-Eijnden, Lenka ZdeborováNeurIPS 2020 · 65 citations
Related papers
- Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data SpectraRoman Worschech, Bernd RosenowICLR 2025
- Scaling Laws in Linear Regression: Compute, Parameters, and DataLicong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett et al.NeurIPS 2024 · 57 citations
- Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvaluesJakob Kramp, Javed Lindner, Moritz HeliasICML 2026 · 2 citations
- A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insightsFabiola Ricci, Claudia Merger, Sebastian GoldtICML 2026
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba et al.ICLR 2026 · 20 citations
