Lune

ICML2024Top-tier venue

Why Do You Grok? A Theoretical Analysis on Grokking Modular Addition

Mohamad Amin Mohamadi, Zhiyuan Li, Lei Wu, Danica J. Sutherland

2024Year
6Top-tier citations

Abstract

We present a theoretical explanation of the "grokking" phenomenon (Power et al., 2022) , where a model generalizes long after overfitting, for the originally-studied problem of modular addition. First, we show that early in gradient descent, when the "kernel regime" approximately holds, no permutation-equivariant model can achieve small population error on modular addition unless it sees at least a constant fraction of all possible data points. Eventually, however, models escape the kernel regime. We show that two-layer quadratic networks that achieve zero training loss with bounded ℓ ∞ norm generalize well with substantially fewer training points, and further show such networks exist and can be found by gradient descent with small ℓ ∞ regularization. We further provide empirical evidence that these networks as well as simple Transformers, leave the kernel regime only after initially overfitting. Taken together, our results strongly support the case for grokking as a consequence of the transition from kernel-like behavior to limiting behavior of gradient descent on deep networks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e7212450-7884-4fc2-a09e-49599ef8cc89

Cited by top-tier papers6

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines