Look Before You Leap: Universal Emergent Mechanism for Retrieval in Language Models
Alexandre Variengien, Eric Winsor
Abstract
When solving challenging problems, language models (LMs) are able to identify relevant information from long and complicated contexts. To study how LMs solve retrieval tasks in diverse situations, we introduce ORION, a collection of structured retrieval tasks, from text understanding to coding. We apply causal analysis on ORION for 18 open-source language models with sizes ranging from 125 million to 70 billion parameters. We find that LMs internally decompose retrieval tasks in a modular way: middle layers at the last token position process the request, while late layers retrieve the correct entity from the context. Building on our high-level understanding, we demonstrate a proof of concept application for scalable internal oversight of LMs to mitigate prompt-injection while requiring human supervision on only a single input.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f4540ff-ff95-409a-a83c-45e8cb399d97Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
Related papers
- Causality-Aided Evaluation and Explanation of Large Language Model-Based Code GenerationZhenlan Ji, Pingchuan Ma, Zongjie Li, Zhaoyu Wang et al.ISSTA 2025 · 1 citation
- METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language ModelsPengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen et al.ACL 2026 · 1 citation
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code UnderstandingAdam Storek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava et al.ACL 2026 · 4 citations
- CogBench: a large language model walks into a psychology labJulian Coda-Forno, Marcel Binz, Jane X. Wang, Eric SchulzICML 2024 · 60 citations
- Causal Detection of Multi-Step LLM Agent AttacksViraaji Mothukuri, Reza M. PariziICML 2026 · 2 citations
