Omnigrok: Grokking Beyond Algorithmic Data
Ziming Liu, Eric J. Michaud, Max Tegmark
Abstract
Grokking, the unusual phenomenon for algorithmic datasets where generalization happens long after overfitting the training data, has remained elusive. We aim to understand grokking by analyzing the loss landscapes of neural networks, identifying the mismatch between training and test loss landscapes as the cause for grokking. We refer to this as the "LU mechanism" because training and test losses (against model weight norm) typically resemble "L" and "U", respectively. This simple mechanism can nicely explain many aspects of grokking: data size dependence, weight decay dependence, the emergence of representations, etc. Guided by the intuitive picture, we are able to induce grokking on tasks involving images, language and molecules. In the reverse direction, we are able to eliminate grokking for algorithmic datasets. We attribute the dramatic nature of grokking for algorithmic datasets to representation learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 181 citations
- Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce GrokkingKaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du et al.ICLR 2024 · 71 citations
- Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of GeneralizationBoshi Wang, Xiang Yue, Yu Su, Huan SunNeurIPS 2024 · 48 citations
- Benign Overfitting and Grokking in ReLU Networks for XOR Cluster DataZhiwei Xu, Yutong Wang, Spencer Frei, Gal Vardi et al.ICLR 2024 · 39 citations
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana et al.NeurIPS 2024 · 39 citations
Builds on7
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Triple descent and the two kinds of overfitting: where & why do they appear?Stéphane d'Ascoli, Levent Sagun, Giulio BiroliNeurIPS 2020 · 94 citations
- Multiple Descent: Design Your Own Generalization CurveLin Chen, Yifei Min, Mikhail Belkin, Amin KarbasiNeurIPS 2021 · 64 citations
Related papers
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- The Geometric Origin of Grokking: Accelerating Generalization via Active Structural ReorganizationKefei Tao, Zhang Zhang, Mingze Qi, Xiaojun DuanICML 2026
- Deep Networks Always Grok and Here is WhyAhmed Imtiaz Humayun, Randall Balestriero, Richard G. BaraniukICML 2024 · 53 citations
- A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationEtienne Boursier, Scott Pesme, Radu-Alexandru DragomirNeurIPS 2025 · 12 citations
