Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
Davide Napolitano, Luca Cagliero, Fabrizio Battiloro
Abstract
The evolution of Visual Large Language Models (VLLMs) has revolutionized the automatic understanding of Visually Rich Documents (VRDs), which contain both textual and visual elements. Although VLLMs excel in Visual Question Answering (VQA) on multi-page VRDs, their ability to detect unanswerable questions is still an open research question. Our research delves into the robustness of the VLLMs to plausible yet unanswerable questions, i.e., questions that appear valid but cannot be answered due to subtle corruptions caused by swaps between related concepts or plausible question formulations. Corruptions are generated by replacing the original natural language entities with other ones of the same type, belonging to different document elements, and in different layout positions or pages of the related document. To this end, we present VRD-UQA (VISUALLY RICH DOCUMENT UNANSWERABLE QUESTION ANSWERING), a benchmark for evaluating VLLMs' resilience to plausible yet unanswerable questions across multiple dimensions. It automatically alters the questions of existing VQA datasets consisting of multi-page VRDs, verifies their unanswerability using a VLLM-as-a-judge approach, and then thoroughly evaluates VLLMs' performance. Experiments, run on 12 models, analyze: (1) The VLLMs' accuracy in detecting unanswerable questions at both page and document levels; (2) The effect of different types of corruption (NLP entity, document element, layout); (3) The effectiveness of different knowledge injection strategies based on in-context learning (OCR, multi-page selection, or the possibility of unanswerability). Our findings reveal VLLMs' limitations and demonstrate that VRD-UQA can serve as an evaluation framework for developing resilient document VQA systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9159498-7f4a-4f7b-87ec-2d24fb55b7deBuilds on6
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz et al.ICCV 2023 · 130 citations
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng et al.CVPR 2024 · 39 citations
- DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma et al.ACL 2024 · 37 citations
- Simulating Errors in Touchscreen TypingDanqing Shi, Yujun Zhu, Francisco Erivaldo Fernandes Junior, Shumin Zhai et al.CHI 2025 · 7 citations
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingChao Deng, Jiale Yuan, Pi Bu, Peijie Wang et al.ACL 2025
Related papers
- Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature RestorationYuxuan Zhou, Baole Wei, Xingjian Hu, Haowei Chen et al.ICML 2026
- Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei et al.ICML 2026 · 1 citation
- Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?Ming Liu, Hao Chen, Jindong Wang, Liwen Wang et al.AAAI 2026
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang et al.CVPR 2026
- Benchmarking Multimodal Large Language Models Against Image CorruptionsXinkuan Qiu, Meina Kan, Yongbin Zhou, Shiguang ShanICCV 2025 · 1 citation
