ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
Duong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le Phuoc
Abstract
In this paper, we propose a new dataset, ReasonVQA, for the Visual Question Answering (VQA) task. Our dataset is automatically integrated with structured encyclopedic knowledge and constructed using a low-cost framework, which is capable of generating complex, multi-hop questions. We evaluated state-of-the-art VQA models on Rea-sonVQA, and the empirical results demonstrate that Rea-sonVQA poses significant challenges to these models, highlighting its potential for benchmarking and advancing the field of VQA. Additionally, our dataset can be easily scaled with respect to input images; the current version surpasses the largest existing datasets requiring external knowledge by more than an order of magnitude. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cbae56b-5134-44a8-9934-5fc7ae046627Cited by top-tier papers2
- Benchmarking PhD-Level Coding in 3D Geometric Computer VisionWenyi Li, Renkai Luo, Yue Yu, Huan-ang Gao et al.CVPR 2026 · 2 citations
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment RewardShizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma et al.CVPR 2026 · 1 citation
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- Encyclopedic VQA: Visual questions about detailed properties of fine-grained categoriesThomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel et al.ICCV 2023 · 111 citations
Related papers
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningKang Chen, Xiangqian WuCVPR 2024
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao et al.ACL 2026
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa et al.AAAI 2023 · 178 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
