From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
Konstantinos C. Tsiolis, Alireza Mousavi-Hosseini, Murat A. Erdogdu
Abstract
To understand feature learning dynamics in neural networks, recent theoretical works have focused on gradient-based learning of Gaussian single-index models, where the label is a nonlinear function of a latent one-dimensional projection of the input. While the sample complexity of online SGD is determined by the information exponent of the link function, recent works improved this by performing multiple gradient steps on the same sample with different learning rates -yielding a non-correlational update rule -and instead are limited by the (potentially much smaller) generative exponent. However, this picture is only valid when these learning rates are sufficiently large. In this paper, we characterize the relationship between learning rate(s) and sample complexity for a broad class of gradient-based algorithms that encapsulates both correlational and non-correlational updates. We demonstrate that, in certain cases, there is a phase transition from an "information exponent regime" with small learning rate to a "generative exponent regime" with large learning rate. Our framework covers prior analyses of one-pass SGD and SGD with batch reuse, while also introducing a new layer-wise training algorithm that leverages a two-timescales approach (via different learning rates for each layer) to go beyond correlational queries without reusing samples or modifying the loss from squared error. Our theoretical study demonstrates that the choice of learning rate is as important as the design of the algorithm in achieving statistical and computational efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd663d35-e819-4e22-9be0-4ead00841c57Builds on26
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 119 citations
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg et al.ICML 2021 · 78 citations
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 77 citations
Related papers
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 · 49 citations
- Online Learning and Information Exponents: The Importance of Batch size & Time/Complexity TradeoffsLuca Arnaboldi, Yatin Dandi, Florent Krzakala, Bruno Loureiro et al.ICML 2024 · 4 citations
- Emergence and scaling laws in SGD learning of shallow neural networksYunwei Ren, Eshaan Nichani, Denny Wu, Jason D. LeeNeurIPS 2025 · 33 citations
- Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index ModelSiyu Chen, Beining Wu, Miao Lu, Zhuoran Yang et al.ICLR 2025
- Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index LearningFilip Kovačević, Hong Chang Ji, Denny Wu, Mahdi Soltanolkotabi et al.ICML 2026 · 2 citations
