A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir
Abstract
We study the dynamics of gradient flow with small weight decay on general training losses . Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay exhibits a two-phase behaviour as . During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of . Then, at time of order , the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the -norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the grokking effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6ed8d50-6fe8-4fd9-a13e-d57cb41e86deCited by top-tier papers2
- Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying RegularizationSebastian Kassing, Simon Weissmann, Leif DöringNeurIPS 2025 · 7 citations
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
Builds on12
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
Related papers
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- Omnigrok: Grokking Beyond Algorithmic DataZiming Liu, Eric J. Michaud, Max TegmarkICLR 2023 · 8 citations
- Grokking at the Edge of Linear SeparabilityAlon Beck, Noam Itzhak Levi, Yohai Bar-SinaiICML 2025
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 1 citation
