Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
Tianyi Lorena Yan, Robin Jia
Abstract
To answer one-to-many factual queries (e.g., listing cities of a country), a language model (LM) must simultaneously recall knowledge and avoid repeating previous answers. How are these two subtasks implemented and integrated internally? Across multiple datasets, models, and prompt templates, we identify a promote-then-suppress mechanism: the model first recalls all answers, and then suppresses previously generated ones. Specifically, LMs use both the subject and previous answer tokens to perform knowledge recall, with attention propagating subject information and MLPs promoting the answers. Then, attention attends to and suppresses previous answer tokens, while MLPs amplify the suppression signal. Our mechanism is corroborated by extensive experimental evidence: in addition to using early decoding and causal tracing, we analyze how components use different tokens by introducing both Token Lens, which decodes aggregated attention updates from specified tokens, and a knockout method that analyzes changes in MLP outputs after removing attention to specified tokens. Overall, we provide new insights into how LMs'internal components interact with different input tokens to support complex factual recall. Code is available at https://github.com/Lorenayannnnn/how-lms-answer-one-to-many-factual-queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Structural Inference: Interpreting Small Language Models with SusceptibilitiesGarrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak et al.ICLR 2026 · 11 citations
- How Language Models Process NegationZhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan MayICML 2026
Builds on14
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 183 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani et al.ICLR 2022 · 74 citations
Related papers
- Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsYifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng et al.EMNLP 2024 · 3 citations
- Characterizing Mechanisms for Factual Recall in Language ModelsQinan Yu, Jack Merullo, Ellie PavlickEMNLP 2023 · 5 citations
- Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language ModelsShaotian Yan, Chen Shen, Wenxiao Wang, Liang Xie et al.ICLR 2025
- Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMsJingcheng Niu, Xingdi Yuan, Tong Wang, Hamidreza Saghir et al.ACL 2025 · 2 citations
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language ModelsMert Yüksekgönül, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar et al.ICLR 2024 · 73 citations
