Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du, Jason D. Lee, Wei Hu
Abstract
Recent work by Power et al. (2022) highlighted a surprising "grokking" phenomenon in learning arithmetic tasks: a neural net first "memorizes" the training set, resulting in perfect training accuracy but near-random test accuracy, and after training for sufficiently longer, it suddenly transitions to perfect test accuracy. This paper studies the grokking phenomenon in theoretical setups and shows that it can be induced by a dichotomy of early and late phase implicit biases. Specifically, when training homogeneous neural nets with large initialization and small weight decay on both classification and regression tasks, we prove that the training process gets trapped at a solution corresponding to a kernel predictor for a long time, and then a very sharp transition to min-norm/max-margin predictors occurs, leading to a dramatic change in test accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f70cb7a-18a2-4de7-a4e2-603adbbdc964Cited by top-tier papers25
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen et al.NeurIPS 2024 · 48 citations
- Uncovering a Universal Abstract Algorithm for Modular Addition in Neural NetworksGavin McCracken, Gabriela Moisescu-Pareja, Vincent Létourneau, Doina Precup et al.NeurIPS 2025 · 14 citations
- Simplicity Bias of Two-Layer Networks beyond Linearly Separable DataNikita Tsoy, Nikola KonstantinovICML 2024 · 12 citations
- A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationEtienne Boursier, Scott Pesme, Radu-Alexandru DragomirNeurIPS 2025 · 12 citations
Builds on27
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
Related papers
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- Grokking in Linear Estimators - A Solvable Model that Groks without UnderstandingNoam Itzhak Levi, Alon Beck, Yohai Bar-SinaiICLR 2024 · 24 citations
- Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker ModelZhiwei Xu, Zhiyu Ni, Yixin Wang, Wei HuICLR 2025
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 43 citations
- Why Do You Grok? A Theoretical Analysis on Grokking Modular AdditionMohamad Amin Mohamadi, Zhiyuan Li, Lei Wu, Danica J. SutherlandICML 2024
