Retrieval Head Mechanistically Explains Long-Context Factuality
Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, Yao Fu
Abstract
Despite the recent progress in long-context large language models (LLMs), it remains elusive how these transformer-based language models acquire the capability to retrieve relevant information from arbitrary locations within the long context. This paper aims to address this question. Our systematic investigation across 4 model families, 6 model scales, and 3 types of finetuning reveals that a special type of attention heads are largely responsible for retrieving relevant information from long context, which we dub retrieval heads. We identify important and intriguing properties of retrieval heads: (1) universal: all the explored models with long-context capability have a set of retrieval heads; (2) sparse: only a small portion (less than 5%) of the attention heads are retrieval. (3) intrinsic: retrieval heads already exist in models pretrained with short context. When extending the context length to 32-128K by continual pretraining, it is still the same set of heads that perform information retrieval. (4) dynamically activated: take Llama-2 7B for example, 12 retrieval heads always attend to the required information no matter how the context is changed. The rest of the retrieval heads are activated in different contexts. (5) causal: completely pruning retrieval heads leads to failure in retrieving relevant information and results in hallucination, while pruning random non-retrieval heads does not affect the model's retrieval ability. We further show that retrieval heads strongly influence chain-of-thought (CoT) reasoning, where the model needs to frequently refer back the question and previously-generated context. Conversely, tasks where the model directly generates the answer using its intrinsic knowledge are less impacted by masking out retrieval heads. These observations collectively explain which internal part of the model seeks information from the input tokens. We believe our insights on retrieval heads foster future research on reducing hallucination, improving reasoning, and compressing the KV cache.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa84df59-0588-4e90-b0ff-dabd794c19fcCited by top-tier papers103
- LoFiT: Localized Fine-tuning on LLM RepresentationsFangcong Yin, Xi Ye, Greg DurrettNeurIPS 2024 · 74 citations
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang et al.NeurIPS 2024 · 71 citations
- Twilight: Adaptive Attention Sparsity with Hierarchical Top- PruningChaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang et al.NeurIPS 2025 · 53 citations
- Mechanistic Detection and Mitigation of Hallucination in Large Reasoning ModelsZhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang et al.ICLR 2026 · 30 citations
- Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsZhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang et al.ACL 2025 · 22 citations
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue et al.ICML 2024 · 204 citations
Related papers
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval AugmentationPatrick Kahardipraja, Reduan Achtibat, Thomas Wiegand, Wojciech Samek et al.NeurIPS 2025 · 13 citations
- Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-rankingWuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen et al.EMNLP 2025
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR HeadsIngeol Baek, Hwan Chang, Sunghyun Ryu, Hwanhee LeeEMNLP 2025
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani et al.NeurIPS 2025 · 5 citations
