Lune

ICML2025Top-tier venue

Grokking Beyond the Euclidean Norm of Model Parameters

Pascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume Rabusseau

2025Year
3Top-tier citations

Abstract

Grokking refers to a delayed generalization following overfitting when optimizing artificial neural networks with gradient-based methods. In this work, we demonstrate that grokking can be induced by regularization, either explicit or implicit. More precisely, we show that when there exists a model with a property P (e.g., sparse or low-rank weights) that generalizes on the problem of interest, gradient descent with a small but non-zero regularization of P (e.g., ℓ 1 or nuclear norm regularization) results in grokking. This extends previous work showing that small nonzero weight decay induces grokking. Moreover, our analysis shows that over-parameterization by adding depth makes it possible to grok or ungrok without explicitly using regularization, which is impossible in shallow cases. We further show that the ℓ 2 norm is not a reliable proxy for generalization when the model is regularized toward a different property P , as the ℓ 2 norm grows in many cases where no weight decay is used, but the model generalizes anyway. We also show that grokking can be amplified solely through data selection, with any other hyperparameter fixed.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0678e859-b5c2-4ee9-8513-19c84749e8a9

Cited by top-tier papers3

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines