Comparing the learning dynamics of in-context learning and fine-tuning in language models
Basile Confavreux, Aaditya K Singh, Jin Hwa Lee, Amaury Sabran, Andrew M Saxe
摘要
Pretrained language models can acquire novel tasks either through in-context learning (ICL)---adapting behavior via activations without weight updates---or through supervised fine-tuning (SFT), where parameters are explicitly updated. Prior work has reported differences in their generalization performance and inductive biases, but the origins of these differences remain poorly understood. In this work, we treat ICL and SFT as distinct learning algorithms and directly compare the learning dynamics they induce across medium-sized models, analyzing both the evolution of their inductive biases and the underlying internal representations. We find that ICL preserves rich input representations but imposes stronger priors inherited from pretraining, whereas SFT suppresses task-irrelevant features---potentially explaining its weaker generalization in few-shot regimes. These results highlight a mechanistic distinction between context-driven and weight-driven learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
相关 Paper
- The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language ModelsDiego Doimo, Alessandro Serra, Alessio Ansuini, Alberto CazzanigaNeurIPS 2024 · 被引用 21 次
- Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight ForgettingSuraj Anand, Michael A. Lepori, Jack Merullo, Ellie PavlickICLR 2025
- Revisiting In-context Learning Inference Circuit in Large Language ModelsHakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya InoueICLR 2025
- IA2: Alignment with ICL Activations improves Supervised Fine-TuningAayush Mishra, Daniel Khashabi, Anqi LiuICLR 2026 · 被引用 1 次
- Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning PerspectiveBishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu 等ACL 2026
