Li2: A Framework on Dynamics of Feature Emergence and Delayed Generalization
Yuandong Tian
摘要
While the phenomenon of grokking, i.e., delayed generalization, has been studied extensively, it remains an open problem whether there is a mathematical framework that characterizes what kind of features will emerge, how and in which conditions it happens, and is still closely connected with the gradient dynamics of the training, for complex structured inputs. We propose a novel framework, named , that captures three key stages for the grokking behavior of 2-layer nonlinear networks: (I) Lazy learning, (II) independent feature learning and (III) interactive feature learning. At the lazy learning stage, top layer overfits to random hidden representation and the model appears to memorize. During lazy learning, the backpropagated gradient from the top layer carries information about the target label, with a specific structure that enables each hidden node to learn their representation independently. Interestingly, the independent dynamics follows exactly the gradient ascent of an energy function , and its local maxima are precisely the emerging features. We study whether these local-optima induced features are generalizable, their representation power, and how they change on sample size, in group arithmetic tasks. When hidden nodes start to interact in the later stage of learning, we provably show how changes to focus on missing features that need to be learned. Our study sheds lights on roles played by key hyperparameters such as weight decay, learning rate and sample sizes in grokking, leads to provable scaling laws of feature emergence, memorization and generalization, and reveals the underlying cause why recent optimizers such as Muon can be effective, from the first principles of gradient dynamics. Our analysis can be extended to multi-layer architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud 等NeurIPS 2022 · 被引用 299 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
相关 Paper
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 被引用 86 次
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 被引用 43 次
- Grokking in Linear Estimators - A Solvable Model that Groks without UnderstandingNoam Itzhak Levi, Alon Beck, Yohai Bar-SinaiICLR 2024 · 被引用 24 次
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
