Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
Tianyu He, Darshil Doshi, Aritra Das, Andrey Gromov
Abstract
Large language models can solve tasks that were not present in the training set. This capability is believed to be due to in-context learning and skill composition. In this work, we study the emergence of in-context learning and skill composition in a collection of modular arithmetic tasks. Specifically, we consider a finite collection of linear modular functions labeled by the vector . We use some of these tasks for pre-training and the rest for out-of-distribution testing. We empirically show that a GPT-style transformer exhibits a transition from in-distribution to out-of-distribution generalization as the number of pre-training tasks increases. We find that the smallest model capable of out-of-distribution generalization requires two transformer blocks, while for deeper models, the out-of-distribution generalization phase is transient, necessitating early stopping. Finally, we perform an interpretability study of the pre-trained models, revealing highly structured representations in both attention heads and MLPs; and discuss the learned algorithms. Notably, we find an algorithmic shift in deeper models, as we go from few to many in-context examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f205305c-224f-4e21-a6c7-45957afb3a91Cited by top-tier papers18
- In-Context Learning Strategies Emerge RationallyDaniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka et al.NeurIPS 2025 · 19 citations
- Uncovering a Universal Abstract Algorithm for Modular Addition in Neural NetworksGavin McCracken, Gabriela Moisescu-Pareja, Vincent Létourneau, Doina Precup et al.NeurIPS 2025 · 14 citations
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 14 citations
- Circuit Stability Characterizes Language Model GeneralizationAlan SunACL 2025 · 4 citations
- Emergent Analogical Reasoning in TransformersGouki Minegishi, Jingyuan Feng, Hiroki Furuta, Takeshi Kojima et al.ICML 2026 · 4 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 324 citations
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud et al.NeurIPS 2022 · 299 citations
Related papers
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
- Task Generalization with Autoregressive Compositional Structure: Can Learning from D Tasks Generalize to DT Tasks?Amirhesam Abedsoltan, Huaqing Zhang, Kaiyue Wen, Hongzhou Lin et al.ICML 2025
- Extrapolation by Association: Length Generalization Transfer In TransformersZiyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak et al.NeurIPS 2025 · 13 citations
- Teaching Arithmetic to Small TransformersNayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee et al.ICLR 2024 · 128 citations
- How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with RepresentationsTianyu Guo, Wei Hu, Song Mei, Huan Wang et al.ICLR 2024 · 80 citations
