Grokking at the Edge of Numerical Stability
Lucas Prieto, Melih Barsbey, Pedro A. M. Mediano, Tolga Birdal
摘要
Grokking, or sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon that has challenged our understanding of deep learning. While a lot of progress has been made in understanding grokking, it is still not clear why generalization is delayed and why grokking often does not happen without regularization. In this work we argue that without regularization, grokking tasks push models to the edge of numerical stability, introducing floating point errors in the Softmax that we refer to as Softmax Collapse (SC). We show that SC prevents grokking and that mitigating SC leads to grokking without regularization. Investigating the root cause of SC, we find that beyond the point of overfitting, the gradients strongly align with what we call the naïve loss minimization (NLM) direction. This component of the gradient does not change the predictions of the model but decreases the loss by scaling the logits, usually through the scaling of the weights along their current direction. We show that this scaling of the logits explains the delay in generalization characteristic of grokking, and eventually leads to SC, stopping learning altogether. To validate these hypotheses, we introduce two key contributions that mitigate the issues faced in grokking tasks: (i) StableMax, a new activation function that prevents SC and enables grokking without regularization, and (ii) ⊥ Grad, a training algorithm that leads to quick generalization in grokking tasks by preventing NLM altogether. These contributions provide new insights into grokking, shedding light on its delayed generalization, reliance on regularization, and the effectiveness of known grokking-inducing methods. Code for this paper can be found at: https://github.com/LucasPrietoAl/ grokking-at-the-edge-of-numerical-stability .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren 等NeurIPS 2025 · 被引用 314 次
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without TestZiyue Li, Chenrui Fan, Tianyi ZhouICLR 2026 · 被引用 11 次
- What Happens During the Loss Plateau? Understanding Abrupt Learning in TransformersPulkit Gopalani, Wei HuNeurIPS 2025 · 被引用 6 次
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -ParityJianhao Huang, Baharan MirzasoleimanICML 2026 · 被引用 2 次
- Intrinsic Task Symmetry Drives Generalization in Algorithmic TasksHyeonbin Hwang, Yeachan ParkICML 2026 · 被引用 1 次
它引用的顶会 Paper20
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
相关 Paper
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
- Grokking Beyond the Euclidean Norm of Model ParametersPascal Tikeng Notsawo Jr., Guillaume Dumas, Guillaume RabusseauICML 2025
- Egalitarian Gradient Descent: A Simple Approach to Accelerated GrokkingAli Saheb Pasand, Elvis DohmatobICLR 2026 · 被引用 1 次
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingTing Han, Linara Adilova, Henning Petzka, Jens Kleesiek 等NeurIPS 2025 · 被引用 9 次
- Grokking at the Edge of Linear SeparabilityAlon Beck, Noam Itzhak Levi, Yohai Bar-SinaiICML 2025
