Emergence and scaling laws in SGD learning of shallow neural networks
Yunwei Ren, Eshaan Nichani, Denny Wu, Jason D. Lee
Abstract
We study the complexity of online stochastic gradient descent (SGD) for learning a two-layer neural network with ๐ neurons on isotropic Gaussian data:
where the activation ๐ is an even function with information exponent ๐ * > 2 (defined as the lowest degree in Hermite expansion), v * ๐ ๐โ [ ๐] โ R ๐ are orthonormal signal directions, and non-negative second-layer coefficients satisfy ๐ ๐ 2 ๐ = 1. We focus on the challenging "extensive-width" regime ๐ โซ 1 and permit diverging condition number in the second-layer, covering as a special case the power-law scaling ๐ ๐ โ ๐ -๐ฝ where ๐ฝ โ R โฅ0 . We provide a precise analysis of SGD dynamics for the training of a student two-layer network to minimize the mean squared error (MSE) objective, and identify sharp transition times to recover each signal direction. In the power-law setting, we characterize scaling law exponents for the MSE loss with respect to the number of training samples and SGD steps, as well as the number of trainable parameters. Our analysis entails that while the learning of individual teacher neurons exhibits abrupt transitions, the juxtaposition of ๐ โซ 1 emergent learning curves at different timescales leads to a smooth scaling law in the cumulative objective.
Remark. We focus on high information exponent IE(๐) > 2 link functions as in [OSSW24, SBH24, GWB25]. This setting entails that the learning of each single-index task is "hard" in the sense that online SGD exhibits a long loss plateau, and we utilize this assumption to prove (approximate) decoupling of individual tasks. The condition on even ๐ simplifies the analysis by removing the 1/2 probability of neurons initialized in the wrong hemisphere (see e.g., [BAGJ21]).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Learning quadratic neural networks in high dimensions: SGD dynamics and scaling lawsGรฉrard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural, Denny WuNeurIPS 2025 ยท 23 citations
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba et al.ICLR 2026 ยท 20 citations
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada et al.NeurIPS 2025 ยท 15 citations
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 ยท 7 citations
- Improved Scaling Laws in Linear Regression via Data ReuseLicong Lin, Jingfeng Wu, Peter L. BartlettNeurIPS 2025 ยท 7 citations
Builds on30
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 ยท 179 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 ยท 173 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 ยท 119 citations
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborovรกNeurIPS 2021 ยท 109 citations
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 ยท 100 citations
Related papers
- Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data SpectraRoman Worschech, Bernd RosenowICLR 2025
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 ยท 49 citations
- Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity TradeoffsLuca Arnaboldi, Yatin Dandi, Florent Krzakala, Bruno Loureiro et al.ICML 2024 ยท 4 citations
- Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent AnalysisYunwei Ren, Jason D. LeeNeurIPS 2025 ยท 7 citations
- From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGDKonstantinos C. Tsiolis, Alireza Mousavi-Hosseini, Murat A. ErdogduNeurIPS 2025 ยท 2 citations
