DocVXQA: Context-Aware Visual Explanations for Document Question Answering
Mohamed Ali Souibgui, Changkyu Choi, Andrey Barsky, Kangsoo Jung, Ernest Valveny, Dimosthenis Karatzas
Abstract
We propose DocVXQA, a novel framework for visually self-explainable document question answering. The framework is designed not only to produce accurate answers to questions but also to learn visual heatmaps that highlight contextually critical regions, thereby offering interpretable justifications for the model's decisions. To integrate explanations into the learning process, we quantitatively formulate explainability principles as explicit learning objectives. Unlike conventional methods that emphasize only the regions pertinent to the answer, our framework delivers explanations that are contextually sufficient while remaining representation-efficient. This fosters user trust while achieving a balance between predictive performance and interpretability in DocVQA applications. Extensive experiments, including human evaluation, provide strong evidence supporting the effectiveness of our method. The code is available at https://github.com/ dali92002/DocVXQA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2caf35f2-43be-465b-b1e4-159193102b58Cited by top-tier papers3
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language ModelsJiyoon Pyo, Yuankun Jiao, Dongwon Jung, Zekun Li et al.ICLR 2026
- A Progressive Evidence Localization Framework Based on Wasserstein Gradient Flows for Document Visual Question AnsweringHaosen Wang, Jing Xiao, Mengqiao Li, Xuanze Wang et al.ICML 2026
- Flow-Based Page Unique Semantic Mapping Architecture for Document Visual Question AnsweringHaosen Wang, Jing Xiao, Chaochao Du, Xiaowang Zhang et al.ACL 2026
Builds on15
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- DocFormerv2: Local Features for Document UnderstandingSrikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran et al.AAAI 2024 · 68 citations
Related papers
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- Equivariant and Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Tat-Seng ChuaACM MM 2022 · 33 citations
- F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question AnsweringHendrik Schuff, Heike Adel, Ngoc Thang VuEMNLP 2020
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li et al.AAAI 2024 · 12 citations
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji et al.CVPR 2022 · 108 citations
