Few-Shot Multimodal Explanation for Visual Question Answering
Dizhan Xue, Shengsheng Qian, Changsheng Xu
Abstract
A key object in eXplainable Artificial Intelligence (XAI) is to create intelligent systems capable of reasoning and explaining real-world data to facilitate reliable decision-making. Recent studies have acknowledged the importance of providing user-friendly and verifiable explanations to facilitate trustworthy Visual Question Answering (VQA) systems. This paper aims to promote explainable VQA from both data and method perspectives. First, we propose a new Standard Multimodal Explanation (SME) dataset and a new Few-Shot Multimodal Explanation for VQA (FS-MEVQA) task, which aims to generate the multimodal explanation of the underlying reasoning process for solving visual questions with few training samples. Our SME dataset includes 1,028,230 samples composed of questions, images, answers, and multimodal explanations, which can facilitate research in both traditional MEVQA and FS-MEVQA. To the best of our knowledge, this is the first large-scale dataset with joint language-vision explanations based on standard English and additional visual grounding tokens. Second, we propose a training-free Multimodal Explaining Agent (MEAgent) method based on an LLM agent with multimodal open-world tools to infer answers and generate multimodal explanations for visual questions. Our MEAgent can learn multimodal explanation from merely N(=16) training samples and leverage open-world abilities to perform FS-MEVQA on test samples. Comprehensive experimental results evaluated by language quality metrics, visual detection metric, and visual attribution metrics on our SME dataset indicate the superiority of our method for FS-MEVQA. Our code and data are available at https://github.com/LivXue/FS-MEVQA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 153c9871-a61d-4e76-9f61-a146f99f346cCited by top-tier papers2
- Analyzing Fine-Tuning Representation Shift for Multimodal LLMs SteeringPegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny et al.ICCV 2025 · 17 citations
- SoMe: A Realistic Benchmark for LLM-based Social Media AgentsDizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu et al.AAAI 2026 · 1 citation
Related papers
- REX: Reasoning-aware and Grounded ExplanationShi Chen, Qi ZhaoCVPR 2022 · 24 citations
- Towards a Multimodal Large Language Model with Pixel-Level Insight for BiomedicineXiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang et al.AAAI 2025 · 32 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question AnsweringZhihao Wen, Wenkang Wei, Yuan Fang, Xingtong Yu et al.CVPR 2026 · 1 citation
