Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu, Adityanarayanan Radhakrishnan, Parthe Pandit, Mikhail Belkin
摘要
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process. It is often taken as an example of "emergence", where model ability manifests sharply through a phase transition. In this work, we show that the phenomenon of grokking is not specific to neural networks nor to gradient descent-based optimization. Specifically, we show that this phenomenon occurs when learning modular arithmetic with Recursive Feature Machines (RFM), an iterative algorithm that uses the Average Gradient Outer Product (AGOP) to enable task-specific feature learning with general machine learning models. When used in conjunction with kernel machines, iterating RFM results in a fast transition from random, near zero, test accuracy to perfect test accuracy. This transition cannot be predicted from the training loss, which is identically zero, nor from the test loss, which remains constant in initial iterations. Instead, as we show, the transition is completely determined by feature learning: RFM gradually learns block-circulant features to solve modular arithmetic. Paralleling the results for RFM, we show that neural networks that solve modular arithmetic also learn block-circulant features. Furthermore, we present theoretical evidence that RFM uses such block-circulant features to implement the Fourier Multiplication Algorithm, which prior work posited as the generalizing solution neural networks learn on these tasks. Our results demonstrate that emergence can result purely from learning task-relevant features and is not specific to neural architectures nor gradient descent-based optimization methods. Furthermore, our work provides more evidence for AGOP as a key mechanism for feature learning in neural networks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular dataDaniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan, Mikhail BelkinICLR 2026 · 被引用 18 次
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada 等NeurIPS 2025 · 被引用 15 次
- Uncovering a Universal Abstract Algorithm for Modular Addition in Neural NetworksGavin McCracken, Gabriela Moisescu-Pareja, Vincent Létourneau, Doina Precup 等NeurIPS 2025 · 被引用 14 次
- The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time AdaptationGuannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia YeCVPR 2026 · 被引用 2 次
- FACT: a first-principles alternative to the Neural Feature Ansatz for how networks learn representationsEnric Boix Adserà, Neil Mallinar, James B. Simon, Misha BelkinICLR 2026 · 被引用 2 次
它引用的顶会 Paper13
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud 等NeurIPS 2022 · 被引用 299 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 被引用 181 次
- Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce GrokkingKaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du 等ICLR 2024 · 被引用 71 次
相关 Paper
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- Grokking Finite-Dimensional AlgebraPascal Jr Tikeng Notsawo, Guillaume Dumas, Guillaume RabusseauICML 2026
- Grokking as a First Order Phase Transition in Two Layer NetworksNoa Rubin, Inbar Seroussi, Zohar RingelICLR 2024 · 被引用 43 次
- Li2: A Framework on Dynamics of Feature Emergence and Delayed GeneralizationYuandong TianICLR 2026
- Why Do You Grok? A Theoretical Analysis on Grokking Modular AdditionMohamad Amin Mohamadi, Zhiyuan Li, Lei Wu, Danica J. SutherlandICML 2024
