Lune

NeurIPS2025Top-tier venue

A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation

Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir

2025Year
12Citations
2Top-tier citations

Abstract

We study the dynamics of gradient flow with small weight decay on general training losses F:Rd→RF: \mathbb{R}^d \to \mathbb{R}. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay λ\lambda exhibits a two-phase behaviour as λ→0\lambda \to 0. During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of FF. Then, at time of order 1/λ1/\lambda, the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the ℓ2\ell_2-norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the grokking effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b6ed8d50-6fe8-4fd9-a13e-d57cb41e86de

Cited by top-tier papers2

Ask how each one uses it

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines