When can in-context learning generalize out of task distribution?
Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. Schwab
Abstract
In-context learning (ICL) is a remarkable capability of pretrained transformers that allows models to generalize to unseen tasks after seeing only a few examples. We investigate empirically the conditions necessary on the pretraining distribution for ICL to emerge and generalize out-ofdistribution. Previous work has focused on the number of distinct tasks necessary in the pretraining dataset. Here, we use a different notion of task diversity to study the emergence of ICL in transformers trained on linear functions. We find that as task diversity increases, transformers undergo a transition from a specialized solution, which exhibits ICL only within the pretraining task distribution, to a solution which generalizes out of distribution to the entire task space. We also investigate the nature of the solutions learned by the transformer on both sides of the transition, and observe similar transitions in nonlinear regression problems. We construct a phase diagram to characterize how our concept of task diversity interacts with the number of pretraining tasks. In addition, we explore how factors such as the depth of the model and the dimensionality of the regression problem influence the transition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78508d26-ff5e-47dc-8caf-4f9ece5d2dfeCited by top-tier papers4
- Pretrain–Test Task Alignment Governs Generalization in In-Context LearningMary Letey, Jacob A Zavatone-Veth, Yue M. Lu, Cengiz PehlevanICLR 2026 · 6 citations
- Generalization vs Specialization under Concept ShiftAlex Nguyen, David J. Schwab, Vudtiwat NgampruetikornNeurIPS 2025 · 3 citations
- The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning AlgorithmsJinghan Zhang, Zerui Cheng, Shiqi Chen, Ge Zhang et al.ICML 2026
- How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-OffWaïss Azizian, Ali HasanICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
Related papers
- Pretraining task diversity and the emergence of non-Bayesian in-context learning for regressionAllan Raventós, Mansheej Paul, Feng Chen, Surya GanguliNeurIPS 2023 · 174 citations
- Differential learning kinetics govern the transition from memorization to generalization during in-context learningAlex Nguyen, Gautam ReddyICLR 2025
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman et al.ICLR 2024 · 94 citations
- Can In-context Learning Really Generalize to Out-of-distribution Tasks?Qixun Wang, Yifei Wang, Xianghua Ying, Yisen WangICLR 2025
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 79 citations
