Characterizing Mechanisms for Factual Recall in Language Models
Qinan Yu, Jack Merullo, Ellie Pavlick
Abstract
Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context. These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict. On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations. Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., “The capital of Poland is London”) to overwrite what it learned in pretraining (“Warsaw”). On Pythia and GPT2, the training frequency of both the query country (”Poland”) and the in-context city (”London”) highly affect the models’ likelihood of using the counterfactual. We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits. By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data. This method can increase the rate of generating the in-context answer to 88% of the time simply by scaling a single head at runtime. Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers46
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationHanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian et al.NeurIPS 2024 · 43 citations
- Mechanistic Detection and Mitigation of Hallucination in Large Reasoning ModelsZhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang et al.ICLR 2026 · 30 citations
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval AugmentationPatrick Kahardipraja, Reduan Achtibat, Thomas Wiegand, Wojciech Samek et al.NeurIPS 2025 · 13 citations
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning ModelsYuhui Wang, Changjiang Li, Guangke Chen, Jiacheng Liang et al.ICLR 2026 · 13 citations
Builds on8
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris et al.ICLR 2023 · 50 citations
Related papers
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language ModelsFrancesco Ortu, Zhijing Jin, Diego Doimo, Alberto CazzanigaACL 2026 · 7 citations
- The Effect of Scaling, Retrieval Augmentation and Form on the Factual Consistency of Language ModelsLovisa Hagström, Denitsa Saynova, Tobias Norlund, Moa Johansson et al.EMNLP 2023 · 6 citations
- How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language ModelsMinsung Kim, Dong-Kyum Kim, Jea Kwon, Nakyeong Yang et al.ACL 2026 · 2 citations
- Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMsJingcheng Niu, Xingdi Yuan, Tong Wang, Hamidreza Saghir et al.ACL 2025 · 2 citations
- Tokens to Types: Context Editing with Selective Entity Abstraction for Grounded GenerationRounak Sharma, Debabrata Mahapatra, Shiv Kumar SainiSIGIR 2026
