Can In-context Learning Really Generalize to Out-of-distribution Tasks?
Qixun Wang, Yifei Wang, Xianghua Ying, Yisen Wang
摘要
In this work, we explore the mechanism of in-context learning (ICL) on out-of-distribution (OOD) tasks that were not encountered during training. To achieve this, we conduct synthetic experiments where the objective is to learn OOD mathematical functions through ICL using a GPT-2 model. We reveal that Transformers may struggle to learn OOD task functions through ICL. Specifically, ICL performance resembles implementing a function within the pretraining hypothesis space and optimizing it with gradient descent based on the in-context examples. Additionally, we investigate ICL's well-documented ability to learn unseen abstract labels in context. We demonstrate that such ability only manifests in the scenarios without distributional shifts and, therefore, may not serve as evidence of new-task-learning ability. Furthermore, we assess ICL's performance on OOD tasks when the model is pretrained on multiple tasks. Both empirical and theoretical analyses demonstrate the existence of the low-test-error preference of ICL, where it tends to implement the pretraining function that yields low test error in the testing context. We validate this through numerical experiments. This new theoretical result, combined with our empirical findings, elucidates the mechanism of ICL in addressing OOD tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Generalization vs Specialization under Concept ShiftAlex Nguyen, David J. Schwab, Vudtiwat NgampruetikornNeurIPS 2025 · 被引用 3 次
- SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without TrainingQi Zhang, Yifei Wang, Xiaohan Wang, Jiajun Chai 等ICLR 2026 · 被引用 1 次
- Toward Robust In-Context Learning: Leveraging Out-of-distribution Proxies for Target Inaccessible Demonstration RetrievalHao Xu, Rite Bo, Fausto Giunchiglia, Yingji Li 等ACL 2026
- Counterfactual-based Cognitive Alignment In-Context Learning for Relation ExtractionQibin Li, Shengyuan Bai, Nai Zhou, Nianmin YaoAAAI 2026
- How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-OffWaïss Azizian, Ali HasanICML 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
相关 Paper
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
- Pretrain–Test Task Alignment Governs Generalization in In-Context LearningMary Letey, Jacob A Zavatone-Veth, Yue M. Lu, Cengiz PehlevanICLR 2026 · 被引用 6 次
- Dual Operating Modes of In-Context LearningZiqian Lin, Kangwook LeeICML 2024 · 被引用 50 次
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 被引用 79 次
- How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with RepresentationsTianyu Guo, Wei Hu, Song Mei, Huan Wang 等ICLR 2024 · 被引用 80 次
