REX: Reasoning-aware and Grounded Explanation
Shi Chen, Qi Zhao
摘要
Effectiveness and interpretability are two essential properties for trustworthy AI systems. Most recent studies in visual reasoning are dedicated to improving the accuracy of predicted answers, and less attention is paid to explaining the rationales behind the decisions. As a result, they commonly take advantage of spurious biases instead of actually reasoning on the visual-textual data, and have yet developed the capability to explain their decision making by considering key information from both modalities. This paper aims to close the gap from three distinct perspectives: first, we define a new type of multi-modal explanations that explain the decisions by progressively traversing the reasoning process and grounding keywords in the images. We develop a functional program to sequentially ex-ecute different reasoning steps and construct a new dataset with 1,040,830 multi-modal explanations. Second, we iden-tify the critical need to tightly couple important components across the visual and textual modalities for explaining the decisions, and propose a novel explanation generation method that explicitly models the pairwise correspon-dence between words and regions of interest. It improves the visual grounding capability by a considerable margin, resulting in enhanced interpretability and reasoning performance. Finally, with our new data and method, we perform extensive analyses to study the effectiveness of our explanation under different settings, including multi-task learning and transfer learning. Our code and data are available at https://github.com/szzexpoi/rex.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 被引用 20 次
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 被引用 19 次
- Analyzing Fine-Tuning Representation Shift for Multimodal LLMs SteeringPegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny 等ICCV 2025 · 被引用 17 次
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li 等AAAI 2024 · 被引用 12 次
- DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision ModelsSimone Carnemolla, Matteo Pennisi, Sarinda Samarasinghe, Giovanni Bellitto 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper4
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami 等NeurIPS 2021 · 被引用 1,020 次
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda 等ICCV 2019 · 被引用 482 次
- Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian WolfCVPR 2021
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris 等CVPR 2021
相关 Paper
- Few-Shot Multimodal Explanation for Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuACM MM 2024 · 被引用 5 次
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu 等EMNLP 2025
- Multi-modal Action Chain Abductive ReasoningMengze Li, Tianbao Wang, Jiahe Xu, Kairong Han 等ACL 2023 · 被引用 11 次
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou 等CVPR 2026 · 被引用 2 次
- R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal FormalizationYi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang 等ICCV 2025 · 被引用 21 次
