Emergence and scaling laws in SGD learning of shallow neural networks
Yunwei Ren, Eshaan Nichani, Denny Wu, Jason D. Lee
摘要
We study the complexity of online stochastic gradient descent (SGD) for learning a two-layer neural network with 𝑃 neurons on isotropic Gaussian data:
where the activation 𝜎 is an even function with information exponent 𝑘 * > 2 (defined as the lowest degree in Hermite expansion), v * 𝑝 𝑝∈ [ 𝑃] ⊂ R 𝑑 are orthonormal signal directions, and non-negative second-layer coefficients satisfy 𝑝 𝑎 2 𝑝 = 1. We focus on the challenging "extensive-width" regime 𝑃 ≫ 1 and permit diverging condition number in the second-layer, covering as a special case the power-law scaling 𝑎 𝑝 ≍ 𝑝 -𝛽 where 𝛽 ∈ R ≥0 . We provide a precise analysis of SGD dynamics for the training of a student two-layer network to minimize the mean squared error (MSE) objective, and identify sharp transition times to recover each signal direction. In the power-law setting, we characterize scaling law exponents for the MSE loss with respect to the number of training samples and SGD steps, as well as the number of trainable parameters. Our analysis entails that while the learning of individual teacher neurons exhibits abrupt transitions, the juxtaposition of 𝑃 ≫ 1 emergent learning curves at different timescales leads to a smooth scaling law in the cumulative objective.
Remark. We focus on high information exponent IE(𝜎) > 2 link functions as in [OSSW24, SBH24, GWB25]. This setting entails that the learning of each single-index task is "hard" in the sense that online SGD exhibits a long loss plateau, and we utilize this assumption to prove (approximate) decoupling of individual tasks. The condition on even 𝜎 simplifies the analysis by removing the 1/2 probability of neurons initialized in the wrong hemisphere (see e.g., [BAGJ21]).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Learning quadratic neural networks in high dimensions: SGD dynamics and scaling lawsGérard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural, Denny WuNeurIPS 2025 · 被引用 23 次
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba 等ICLR 2026 · 被引用 20 次
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada 等NeurIPS 2025 · 被引用 15 次
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 · 被引用 7 次
- Improved Scaling Laws in Linear Regression via Data ReuseLicong Lin, Jingfeng Wu, Peter L. BartlettNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper30
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 被引用 119 次
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 被引用 109 次
- Task-Specific Skill Localization in Fine-tuned Language ModelsAbhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, Sanjeev AroraICML 2023 · 被引用 100 次
相关 Paper
- Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data SpectraRoman Worschech, Bernd RosenowICLR 2025
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 49 次
- Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity TradeoffsLuca Arnaboldi, Yatin Dandi, Florent Krzakala, Bruno Loureiro 等ICML 2024 · 被引用 4 次
- Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent AnalysisYunwei Ren, Jason D. LeeNeurIPS 2025 · 被引用 7 次
- From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGDKonstantinos C. Tsiolis, Alireza Mousavi-Hosseini, Murat A. ErdogduNeurIPS 2025 · 被引用 2 次
