Grokking Beyond the Euclidean Norm of Model Parameters
Pascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume Rabusseau
摘要
Grokking refers to a delayed generalization following overfitting when optimizing artificial neural networks with gradient-based methods. In this work, we demonstrate that grokking can be induced by regularization, either explicit or implicit. More precisely, we show that when there exists a model with a property P (e.g., sparse or low-rank weights) that generalizes on the problem of interest, gradient descent with a small but non-zero regularization of P (e.g., ℓ 1 or nuclear norm regularization) results in grokking. This extends previous work showing that small nonzero weight decay induces grokking. Moreover, our analysis shows that over-parameterization by adding depth makes it possible to grok or ungrok without explicitly using regularization, which is impossible in shallow cases. We further show that the ℓ 2 norm is not a reliable proxy for generalization when the model is regularized toward a different property P , as the ℓ 2 norm grows in many cases where no weight decay is used, but the model generalizes anyway. We also show that grokking can be amplified solely through data selection, with any other hyperparameter fixed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Intrinsic Task Symmetry Drives Generalization in Algorithmic TasksHyeonbin Hwang, Yeachan ParkICML 2026 · 被引用 1 次
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 被引用 1 次
- Grokking Finite-Dimensional AlgebraPascal Jr Tikeng Notsawo, Guillaume Dumas, Guillaume RabusseauICML 2026
它引用的顶会 Paper26
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
相关 Paper
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
- Omnigrok: Grokking Beyond Algorithmic DataZiming Liu, Eric J. Michaud, Max TegmarkICLR 2023 · 被引用 8 次
- Why Do You Grok? A Theoretical Analysis on Grokking Modular AdditionMohamad Amin Mohamadi, Zhiyuan Li, Lei Wu, Danica J. SutherlandICML 2024
- On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning RegimeShuai Jiang, Eric C. Cyr, Ben S. Southworth, Alexey VoroninICLR 2026 · 被引用 1 次
- A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationEtienne Boursier, Scott Pesme, Radu-Alexandru DragomirNeurIPS 2025 · 被引用 12 次
