Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
Yuxin Dong, Jiachen Jiang, Zhihui Zhu, Xia Ning
Abstract
Task vector is a compelling mechanism for accelerating inference in in-context learning (ICL) by distilling task-specific information into a single, reusable representation. Despite their empirical success, the underlying principles governing their emergence and functionality remain unclear. This work proposes the Task Vectors as Representative Demonstrations conjecture, positing that task vectors encode single in-context demonstrations distilled from the original ones. We provide both theoretical and empirical support for this conjecture. First, we show that task vectors naturally emerge in linear transformers trained on triplet-formatted prompts through loss landscape analysis. Next, we predict the failure of task vectors in representing high-rank mappings and confirm this on practical LLMs. Our findings are further validated through saliency analyses and parameter visualization, suggesting an enhancement of task vectors by injecting multiple ones into few-shot prompts. Together, our results advance the understanding of task vectors and shed light on the mechanisms underlying ICL in transformer-based models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90fa2480-1ffd-47b6-96e0-0741e7daff59Cited by top-tier papers4
- Mechanism of Task-oriented Information Removal in In-context LearningHakaze Cho, Haolin Yang, Gouki Minegishi, Naoya InoueICLR 2026 · 3 citations
- How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context LearningEntang Wang, Yiwei Wang, Aleksandra Bakalova, Michael HahnICML 2026 · 1 citation
- Beyond Plain Demos: A Demo-Centric Anchoring Paradigm for In-Context Learning in Alzheimer's Disease DetectionPuzhen Su, Haoran Yin, Yongzhu Miao, Jintao Tang et al.AAAI 2026
- Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic InsightsHaolin Yang, Hakaze Cho, Kaize Ding, Naoya InoueICLR 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Do different prompting methods yield a common task representation in language models?Guy Davidson, Todd M. Gureckis, Brenden M. Lake, Adina WilliamsNeurIPS 2025 · 11 citations
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller et al.ICLR 2024 · 229 citations
- Revisiting In-context Learning Inference Circuit in Large Language ModelsHakaze Cho, Mariko Kato, Yoshihiro Sakai, Naoya InoueICLR 2025
- Which Attention Heads Matter for In-Context Learning?Kayo Yin, Jacob SteinhardtICML 2025
- Take Off the Training Wheels! Progressive In-Context Learning for Effective AlignmentZhenyu Liu, Dongfang Li, Xinshuo Hu, Xinping Zhao et al.EMNLP 2024 · 1 citation
