Grokking in Linear Estimators - A Solvable Model that Groks without Understanding
Noam Itzhak Levi, Alon Beck, Yohai Bar-Sinai
Abstract
Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data. We show both analytically and numerically that grokking can surprisingly occur in linear networks performing linear tasks in a simple teacher-student setup with Gaussian inputs. In this setting, the full training dynamics is derived in terms of the training and generalization data covariance matrix. We present exact predictions on how the grokking time depends on input and output dimensionality, train sample size, regularization, and network initialization. We demonstrate that the sharp increase in generalization accuracy may not imply a transition from "memorization" to "understanding", but can simply be an artifact of the accuracy measure. We provide empirical verification for our calculations, along with preliminary results indicating that some predictions also hold for deeper networks, with non-linear activations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 618b4cd6-13d1-474a-a88b-b7b45401e557Cited by top-tier papers10
- Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce GrokkingKaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du et al.ICLR 2024 · 71 citations
- Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondAlan Jeffares, Alicia Curth, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Explaining Grokking and Information Bottleneck through Neural Collapse EmergenceKeitaro Sakamoto, Issei SatoICLR 2026 · 5 citations
- Intrinsic Task Symmetry Drives Generalization in Algorithmic TasksHyeonbin Hwang, Yeachan ParkICML 2026 · 1 citation
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 1 citation
Builds on4
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 144 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
Related papers
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 43 citations
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
- Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker ModelZhiwei Xu, Zhiyu Ni, Yixin Wang, Wei HuICLR 2025
- Li2: A Framework on Dynamics of Feature Emergence and Delayed GeneralizationYuandong TianICLR 2026
