The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
Patrick Kahardipraja, Reduan Achtibat, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin
Abstract
Large language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation. While promising, its inner workings remain unclear. In this work, we shed light on the mechanism of in-context retrieval augmentation for question answering by viewing a prompt as a composition of informational components. We propose an attributionbased method to identify specialized attention heads, revealing in-context heads that comprehend instructions and retrieve relevant contextual information, and parametric heads that store entities' relational knowledge. To better understand their roles, we extract function vectors and modify their attention weights to show how they can influence the answer generation process. Finally, we leverage the gained insights to trace the sources of knowledge used during inference, paving the way towards more safe and transparent language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f52f27a-a5e2-47f1-8c6f-b02e3bd6cbbeCited by top-tier papers7
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context LearningHaolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya InoueNeurIPS 2025 · 11 citations
- Large Vision-Language Models Get Lost in AttentionGongli Xi, Ye Tian, Mengyu Yang, Huahui Yi et al.ICML 2026 · 4 citations
- Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head AnalysisHaolin Yang, Hakaze Cho, Naoya InoueICLR 2026 · 2 citations
- Attribution-Guided DecodingPiotr Komorowski, Elena Golimblevskaia, Reduan Achtibat, Thomas Wiegand et al.ICLR 2026 · 1 citation
- Provable In-Context Vector Arithmetic via Retrieving Task ConceptsDake Bu, Wei Huang, Andi Han, Atsushi Nitanda et al.ICML 2025
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- Retrieval Head Mechanistically Explains Long-Context FactualityWenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng et al.ICLR 2025
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani et al.NeurIPS 2025 · 5 citations
- Retrieval Heads are DynamicYuping Lin, Zitao Li, Yue Xing, Pengfei He et al.ACL 2026
- Understanding Parametric and Contextual Knowledge Reconciliation within Large Language ModelsJun Zhao, Yongzhuo Yang, Xiang Hu, Jingqi Tong et al.NeurIPS 2025 · 10 citations
- Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationJirui Qi, Gabriele Sarti, Raquel Fernández, Arianna BisazzaEMNLP 2024 · 6 citations
