Grokking at the Edge of Numerical Stability
Lucas Prieto, Melih Barsbey, Pedro A. M. Mediano, Tolga Birdal
Abstract
Grokking, or sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon that has challenged our understanding of deep learning. While a lot of progress has been made in understanding grokking, it is still not clear why generalization is delayed and why grokking often does not happen without regularization. In this work we argue that without regularization, grokking tasks push models to the edge of numerical stability, introducing floating point errors in the Softmax that we refer to as Softmax Collapse (SC). We show that SC prevents grokking and that mitigating SC leads to grokking without regularization. Investigating the root cause of SC, we find that beyond the point of overfitting, the gradients strongly align with what we call the naïve loss minimization (NLM) direction. This component of the gradient does not change the predictions of the model but decreases the loss by scaling the logits, usually through the scaling of the weights along their current direction. We show that this scaling of the logits explains the delay in generalization characteristic of grokking, and eventually leads to SC, stopping learning altogether. To validate these hypotheses, we introduce two key contributions that mitigate the issues faced in grokking tasks: (i) StableMax, a new activation function that prevents SC and enables grokking without regularization, and (ii) ⊥ Grad, a training algorithm that leads to quick generalization in grokking tasks by preventing NLM altogether. These contributions provide new insights into grokking, shedding light on its delayed generalization, reliance on regularization, and the effectiveness of known grokking-inducing methods. Code for this paper can be found at: https://github.com/LucasPrietoAl/ grokking-at-the-edge-of-numerical-stability .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b0663d0-0f8c-4c86-895e-e4d81c75e69fCited by top-tier papers10
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without TestZiyue Li, Chenrui Fan, Tianyi ZhouICLR 2026 · 11 citations
- What Happens During the Loss Plateau? Understanding Abrupt Learning in TransformersPulkit Gopalani, Wei HuNeurIPS 2025 · 6 citations
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -ParityJianhao Huang, Baharan MirzasoleimanICML 2026 · 2 citations
- Intrinsic Task Symmetry Drives Generalization in Algorithmic TasksHyeonbin Hwang, Yeachan ParkICML 2026 · 1 citation
Builds on20
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han et al.ICLR 2021 · 165 citations
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 144 citations
Related papers
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 1 citation
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek et al.NeurIPS 2025 · 9 citations
- Grokking at the Edge of Linear SeparabilityAlon Beck, Noam Itzhak Levi, Yohai Bar-SinaiICML 2025
