Diagnosing Evidence Utilization in Multimodal Document Question Answering
Debolena Basak, Digbalay Bose, Koustava Goswami, Maunendra Sankar Desarkar
Abstract
Recent Multimodal Large Language Models (MLLMs) support retrieval-augmented generation (RAG) for document question answering (QA), yet it remains unclear how effectively they use the provided evidence during answer generation. In this work, we conduct a controlled empirical study of 7 popular MLLMs on long multimodal multi-document question answering in a RAG setting. Our analysis shows that zero-shot performance varies substantially across evidence types (e.g., text, image, table, chart, and cross-evidence), revealing a strong reliance on text-based evidence and weaker performance on image-only and cross-evidence inputs. We further find that supervised finetuning yields limited and dataset-dependent improvements, often preserving existing evidence-type disparities. To better understand these behaviours, we perform attention-based analysis to examine how models allocate attention across different token types (image evidence, text evidence, system prompt, question, and output tokens), and find that low attention allocation to image tokens is associated with weaker performance on image-based evidence. Our findings provide insights into how effectively current MLLMs use multimodal evidence in document QA and highlight the key limitations in multimodal document understanding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0088d790-e158-48ed-bf06-06f3982c9ca1Related papers
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question AnsweringAlberto Compagnoni, Marco Morini, Sara Sarto, Federico Cocchi et al.CVPR 2026 · 11 citations
- LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question AnsweringQingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha et al.EMNLP 2024 · 13 citations
- SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document UnderstandingJian Chen, Ruiyi Zhang, Yufan Zhou, Tong Yu et al.ICLR 2025
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun et al.AAAI 2026 · 2 citations
- Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsZhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang et al.ACM MM 2025 · 4 citations
