VISA: Retrieval Augmented Generation with Visual Source Attribution
Xueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Wenhu Chen, Jimmy Lin
Abstract
Generation with source attribution is important for enhancing the verifiability of retrieval-augmented generation (RAG) systems. However, existing approaches in RAG primarily link generated content to document-level references, making it challenging for users to locate evidence among multiple content-rich retrieved documents. To address this challenge, we propose Retrieval-Augmented Generation with Visual Source Attribution (VISA), a novel approach that combines answer generation with visual source attribution. Leveraging large vision-language models (VLMs), VISA identifies the evidence and highlights the exact regions that support the generated answers with bounding boxes in the retrieved document screenshots. To evaluate its effectiveness, we curated two datasets: Wiki-VISA, based on crawled Wikipedia webpage screenshots, and Paper-VISA, derived from PubLayNet and tailored to the medical domain. Experimental results demonstrate the effectiveness of VISA for visual source attribution on documents' original look, as well as highlighting the challenges for improvement. Code, data, and model checkpoints will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language ModelsTobias Schreieder, Tim Schopf, Michael FärberACL 2026 · 10 citations
- DocLens: A Tool-Augmented Multi-Agent Framework for Long Visual Document UnderstandingDawei Zhu, Rui Meng, Jiefeng Chen, Sujian Li et al.ACL 2026 · 10 citations
- Pixel Reasoner: Incentivizing Pixel Space Reasoning via Curiosity-Driven Reinforcement LearningAlex Su, Haozhe Wang, Weiming Ren, Fangzhen Lin et al.NeurIPS 2025 · 6 citations
- Document Screenshot Retrievers are Vulnerable to Pixel Poisoning AttacksShengyao Zhuang, Ekaterina Khramtsova, Xueguang Ma, Bevan Koopman et al.SIGIR 2025 · 4 citations
- Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement LearningShuochen Liu, Pengfei Luo, Chao Zhang, Yuhao Chen et al.AAAI 2026 · 2 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 152 citations
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsOded Ovadia, Menachem Brief, Moshik Mishaeli, Oren ElishaEMNLP 2024 · 89 citations
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsKanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu et al.ACL 2024 · 33 citations
Related papers
- MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question AnsweringXuze Li, Haozhao Wang, Zhenyu Huang, Zhongxu Wang et al.CVPR 2026
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui et al.ICLR 2025
- Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented GenerationJirui Qi, Gabriele Sarti, Raquel Fernández, Arianna BisazzaEMNLP 2024 · 6 citations
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida et al.CVPR 2025
- MIRA: A Novel Framework for Fusing Modalities in Medical RAGJinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen et al.ACM MM 2025 · 5 citations
