The Three Stages of Learning Dynamics in High-dimensional Kernel Methods
Nikhil Ghosh, Song Mei, Bin Yu
摘要
To understand how deep learning works, it is crucial to understand the training dynamics of neural networks. Several interesting hypotheses about these dynamics have been made based on empirically observed phenomena, but there exists a limited theoretical understanding of when and why such phenomena occur. In this paper, we consider the training dynamics of gradient flow on kernel least-squares objectives, which is a limiting dynamics of SGD trained neural networks. Using precise high-dimensional asymptotics, we characterize the dynamics of the fitted model in two"worlds": in the Oracle World the model is trained on the population distribution and in the Empirical World the model is trained on a sampled dataset. We show that under mild conditions on the kernel and target regression function the training dynamics undergo three stages characterized by the behaviors of the models in the two worlds. Our theoretical results also mathematically formalize some interesting deep learning phenomena. Specifically, in our setting we show that SGD progressively learns more complex functions and that there is a"deep bootstrap"phenomenon: during the second stage, the test error of both worlds remain close despite the empirical training error being much smaller. Finally, we give a concrete example comparing the dynamics of two different kernels which shows that faster training is not necessary for better generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 被引用 15 次
- Identifying good directions to escape the NTK regime and efficiently learn low-degree plus sparse polynomialsEshaan Nichani, Yu Bai, Jason D. LeeNeurIPS 2022 · 被引用 15 次
- Beyond Implicit Bias: The Insignificance of SGD Noise in Online LearningNikhil Vyas, Depen Morwani, Rosie Zhao, Gal Kaplun 等ICML 2024 · 被引用 8 次
- High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit RegularizationYihang Chen, Fanghui Liu, Taiji Suzuki, Volkan CevherICML 2024 · 被引用 5 次
它引用的顶会 Paper4
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
- The Deep Bootstrap Framework: Good Online Learners are Good Offline GeneralizersPreetum Nakkiran, Behnam Neyshabur, Hanie SedghiICLR 2021 · 被引用 75 次
相关 Paper
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 被引用 86 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs 等ICML 2020 · 被引用 229 次
- A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationEtienne Boursier, Scott Pesme, Radu-Alexandru DragomirNeurIPS 2025 · 被引用 12 次
- Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter TransferBlake Bordelon, Cengiz PehlevanICML 2025
