Toward Understanding In-context vs. In-weight Learning
Bryan Chan, Xinyi Chen, András György, Dale Schuurmans
摘要
It has recently been demonstrated empirically that in-context learning emerges in transformers when certain distributional properties are present in the training data, but this ability can also diminish upon further training. We provide a new theoretical understanding of these phenomena by identifying simplified distributional properties that give rise to the emergence and eventual disappearance of in-context learning. We do so by first analyzing a simplified model that uses a gating mechanism to choose between an in-weight and an in-context predictor. Through a combination of a generalization error and regret analysis we identify conditions where in-context and in-weight learning emerge. These theoretical findings are then corroborated experimentally by comparing the behaviour of a full transformer on the simplified distributions to that of the stylized model, demonstrating aligned results. We then extend the study to a full large language model, showing how fine-tuning on various collections of natural language prompts can elicit similar in-context and in-weight learning behaviour. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- In-Context Learning Strategies Emerge RationallyDaniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka 等NeurIPS 2025 · 被引用 19 次
- Understanding Prompt Tuning and In-Context Learning via Meta-LearningTim Genewein, Kevin Li, Jordi Grau-Moya, Anian Ruoss 等NeurIPS 2025 · 被引用 10 次
- Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized WorldsFan Wang, Pengtao Shao, Yiming Zhang, Bo Yu 等NeurIPS 2025 · 被引用 6 次
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng 等ICLR 2026 · 被引用 5 次
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 被引用 3 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
相关 Paper
- The Transient Nature of Emergent In-Context Learning in TransformersAaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz, Erin Grant 等NeurIPS 2023 · 被引用 92 次
- Strategy Coopetition Explains the Emergence and Transience of In-Context LearningAaditya K. Singh, Ted Moskovitz, Sara Dragutinovic, Felix Hill 等ICML 2025
- The mechanistic basis of data dependence and abrupt learning in an in-context classification taskGautam ReddyICLR 2024 · 被引用 112 次
- Differential learning kinetics govern the transition from memorization to generalization during in-context learningAlex Nguyen, Gautam ReddyICLR 2025
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
