VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning
Kang Chen, Xiangqian Wu
Abstract
Achieving the optimal form of Visual Question Answering mandates a profound grasp of understanding, grounding, and reasoning within the intersecting domains of vision and language. Traditional VQA benchmarks have predom-inantly focused on simplistic tasks such as counting, visual attributes, and object detection, which do not necessitate intricate cross-modal information understanding and inference. Motivated by the need for a more comprehensive evaluation, we introduce a novel dataset comprising 23,781 questions derived from 10,124 image-text pairs. Specifically, the task of this dataset requires the model to align multimedia representations of the same entity to implement multi-hop reasoning between image and text and finally use natural language to answer the question. Furthermore, we evaluate this VTQA dataset, comparing the performance of both state-of-the-art VQA models and our proposed base-line model, the Key Entity Cross-Media Reasoning Network (KECMRN). The VTQA task poses formidable challenges for traditional VQA models, underscoring its intrinsic complexity. Conversely, KECMRN exhibits a modest improvement, signifying its potential in multimedia entity alignment and multi-step reasoning. Our analysis underscores the diversity, difficulty, and scale of the VTQA task compared to previous multimodal QA datasets. In conclusion, we anticipate that this dataset will serve as a pivotal resource for advancing and evaluating models proficient in multime-dia entity alignment, multi-step reasoning, and open-ended answer generation. Our dataset and code is available at https://visual-text-qa.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72b030d1-9e27-479a-b951-9978673de99fCited by top-tier papers8
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following BehaviorWeikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li et al.CVPR 2026 · 6 citations
- From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic ReasoningYuhui Zeng, Haoxiang Wu, Wenjie Nie, Guangyao Chen et al.ICCV 2025 · 2 citations
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music ReasoningGagan Mundada, Yash Vishe, Amit Namburi, Xin Xu et al.EMNLP 2025 · 1 citation
- RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and ManipulationLinfei Li, Lin Zhang, Ying ShenCVPR 2026
- Attention Grounded Enhancement for Visual Document RetrievalWanqing Cui, Wei Huang, Yazhi Guo, Yibo Hu et al.SIGIR 2026
Builds on1
Related papers
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An et al.CVPR 2026 · 2 citations
- ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringDuong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le PhuocICCV 2025 · 2 citations
- MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingRevanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin et al.AAAI 2022 · 37 citations
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao et al.ACL 2026
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
