The Optimization Landscape of SGD Across the Feature Learning Strength
Alexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz Pehlevan
Abstract
We consider neural networks (NNs) where the final layer is down-scaled by a fixed hyperparameter . Recent work has identified as controlling the strength of feature learning. As increases, network evolution changes from "lazy" kernel dynamics to "rich" feature-learning dynamics, with a host of associated benefits including improved performance on common tasks. In this work, we conduct a thorough empirical investigation of the effect of scaling across a variety of models and datasets in the online training setting. We first examine the interaction of with the learning rate , identifying several scaling regimes in the - plane which we explain theoretically using a simple model. We find that the optimal learning rate scales non-trivially with . In particular, when and when for a feed-forward network of depth . Using this optimal learning rate scaling, we proceed with an empirical study of the under-explored "ultra-rich" regime. We find that networks in this regime display characteristic loss curves, starting with a long plateau followed by a drop-off, sometimes followed by one or more additional staircase steps. We find networks of different large values optimize along similar trajectories up to a reparameterization of time. We further find that optimal online performance is often found at large and could be missed if this hyperparameter is not tuned. Our findings indicate that analytical study of the large- limit may yield useful insights into the dynamics of representation learning in performant models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7a61319-dc50-4bbb-9a9c-73a12e56114eCited by top-tier papers9
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson et al.NeurIPS 2025 · 46 citations
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada et al.NeurIPS 2025 · 15 citations
- Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like ModelsDhruva Karkada, James B. Simon, Yasaman Bahri, Michael R. DeWeeseNeurIPS 2025 · 9 citations
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 · 7 citations
- Predicting Kernel Regression Learning Curves from Only Raw Data StatisticsDhruva Karkada, Joseph Turnbull, Yuxi Liu, James B SimonICLR 2026 · 6 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
Related papers
- From Lazy to Rich: Exact Learning Dynamics in Deep Linear NetworksClémentine Carla Juliette Dominé, Nicolas Anguita, Alexandra Maria Proca, Lukas Braun et al.ICLR 2025
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 84 citations
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann et al.ICML 2025
