The Optimization Landscape of SGD Across the Feature Learning Strength
Alexander B. Atanasov, Alexandru Meterez, James B. Simon, Cengiz Pehlevan
摘要
We consider neural networks (NNs) where the final layer is down-scaled by a fixed hyperparameter . Recent work has identified as controlling the strength of feature learning. As increases, network evolution changes from "lazy" kernel dynamics to "rich" feature-learning dynamics, with a host of associated benefits including improved performance on common tasks. In this work, we conduct a thorough empirical investigation of the effect of scaling across a variety of models and datasets in the online training setting. We first examine the interaction of with the learning rate , identifying several scaling regimes in the - plane which we explain theoretically using a simple model. We find that the optimal learning rate scales non-trivially with . In particular, when and when for a feed-forward network of depth . Using this optimal learning rate scaling, we proceed with an empirical study of the under-explored "ultra-rich" regime. We find that networks in this regime display characteristic loss curves, starting with a long plateau followed by a drop-off, sometimes followed by one or more additional staircase steps. We find networks of different large values optimize along similar trajectories up to a reparameterization of time. We further find that optimal online performance is often found at large and could be missed if this hyperparameter is not tuned. Our findings indicate that analytical study of the large- limit may yield useful insights into the dynamics of representation learning in performant models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson 等NeurIPS 2025 · 被引用 46 次
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada 等NeurIPS 2025 · 被引用 15 次
- Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like ModelsDhruva Karkada, James B. Simon, Yasaman Bahri, Michael R. DeWeeseNeurIPS 2025 · 被引用 9 次
- On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingMoritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena Chennuru VankadaraNeurIPS 2025 · 被引用 7 次
- Predicting Kernel Regression Learning Curves from Only Raw Data StatisticsDhruva Karkada, Joseph Turnbull, Yuxi Liu, James B SimonICLR 2026 · 被引用 6 次
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
相关 Paper
- From Lazy to Rich: Exact Learning Dynamics in Deep Linear NetworksClémentine Carla Juliette Dominé, Nicolas Anguita, Alexandra Maria Proca, Lukas Braun 等ICLR 2025
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- The Importance of Being Lazy: Scaling Limits of Continual LearningJacopo Graldi, Alessandro Breccia, Giulia Lanzillotta, Thomas Hofmann 等ICML 2025
