The Transient Nature of Emergent In-Context Learning in Transformers
Aaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz, Erin Grant, Andrew M. Saxe, Felix Hill
Abstract
Transformer neural networks can exhibit a surprising capacity for in-context learning (ICL) despite not being explicitly trained for it. Prior work has provided a deeper understanding of how ICL emerges in transformers, e.g., through the lens of mechanistic interpretability, Bayesian inference, or by examining the distributional properties of training data. However, in each of these cases, ICL is treated largely as a persistent phenomenon; namely, once ICL emerges, it is assumed to persist asymptotically. Here, we show that the emergence of ICL during transformer training is, in fact, often transient. We train transformers on synthetic data designed so that both ICL and in-weights learning (IWL) strategies can lead to correct predictions. We find that ICL first emerges, then disappears and gives way to IWL, all while the training loss decreases, indicating an asymptotic preference for IWL. The transient nature of ICL is observed in transformers across a range of model sizes and datasets, raising the question of how much to "overtrain" transformers when seeking compact, cheaper-to-run models. We find that L2 regularization may offer a path to more persistent ICL that removes the need for early stopping based on ICL-style validation tasks. Finally, we present initial evidence that ICL transience may be caused by competition between ICL and IWL circuits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f866ffa-6f22-418f-a815-b25dee470eddCited by top-tier papers38
- The mechanistic basis of data dependence and abrupt learning in an in-context classification taskGautam ReddyICLR 2024 · 112 citations
- What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formationAaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan et al.ICML 2024 · 77 citations
- Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasksTianyu He, Darshil Doshi, Aritra Das, Andrey GromovNeurIPS 2024 · 52 citations
- Is In-Context Learning in Large Language Models Bayesian? A Martingale PerspectiveFabian Falck, Ziyu Wang, Christopher C. HolmesICML 2024 · 46 citations
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Strategy Coopetition Explains the Emergence and Transience of In-Context LearningAaditya K. Singh, Ted Moskovitz, Sara Dragutinovic, Felix Hill et al.ICML 2025
- Toward Understanding In-context vs. In-weight LearningBryan Chan, Xinyi Chen, András György, Dale SchuurmansICLR 2025
- Differential learning kinetics govern the transition from memorization to generalization during in-context learningAlex Nguyen, Gautam ReddyICLR 2025
- In-Context Learning Strategies Emerge RationallyDaniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka et al.NeurIPS 2025 · 19 citations
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
