Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
Jirui Qi, Gabriele Sarti, Raquel Fernández, Arianna Bisazza
Abstract
Ensuring the verifiability of model answers is a fundamental challenge for retrieval-augmented generation (RAG) in the question answering (QA) domain. Recently, self-citation prompting was proposed to make large language models (LLMs) generate citations to supporting documents along with their answers. However, self-citing LLMs often struggle to match the required format, refer to non-existent sources, and fail to faithfully reflect LLMs' context usage throughout the generation. In this work, we present MIRAGE -Model Internals-based RAG Explanations -a plug-and-play approach using model internals for faithful answer attribution in RAG applications. MIRAGE detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction via saliency methods. We evaluate our proposed approach on a multilingual extractive QA dataset, finding high agreement with human answer attribution. On open-ended QA, MIRAGE achieves citation quality and efficiency comparable to self-citation while also allowing for a finer-grained control of attribution parameters. Our qualitative evaluation highlights the faithfulness of MIRAGE's attributions and underscores the promising application of model internals for RAG answer attribution. 1 * Equal contribution. 1 Code and data released at https://github.com/ Betswish/MIRAGE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c31b967-09fd-4289-ab63-3402ecec1a91Cited by top-tier papers11
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 118 citations
- Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen et al.ACL 2025 · 64 citations
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval AugmentationPatrick Kahardipraja, Reduan Achtibat, Thomas Wiegand, Wojciech Samek et al.NeurIPS 2025 · 13 citations
- Attributing Response to Context: A Jensen–Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented GenerationRuizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao et al.ICLR 2026 · 11 citations
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 10 citations
Builds on18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- VISA: Retrieval Augmented Generation with Visual Source AttributionXueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon et al.ACL 2025 · 24 citations
- Why Retrieval-Augmented Generation Fails: A Graph PerspectiveKai Guo, Xinnan Dai, Zhibo Zhang, Nuohan Lin et al.KDD 2026
- Faithful in Steps: Improving Generalization and Citation in RAG via Query DecompositionYue Liu, Zhongying Ru, Shimin Di, Jipeng Zhang et al.AAAI 2026
- Learning to Plan and Generate Text with CitationsConstanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao et al.ACL 2024 · 7 citations
- Toward Faithful Retrieval-Augmented Generation with Sparse AutoencodersGuangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha et al.ICLR 2026 · 8 citations
