VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media Reasoning
Kang Chen, Xiangqian Wu
摘要
Achieving the optimal form of Visual Question Answering mandates a profound grasp of understanding, grounding, and reasoning within the intersecting domains of vision and language. Traditional VQA benchmarks have predom-inantly focused on simplistic tasks such as counting, visual attributes, and object detection, which do not necessitate intricate cross-modal information understanding and inference. Motivated by the need for a more comprehensive evaluation, we introduce a novel dataset comprising 23,781 questions derived from 10,124 image-text pairs. Specifically, the task of this dataset requires the model to align multimedia representations of the same entity to implement multi-hop reasoning between image and text and finally use natural language to answer the question. Furthermore, we evaluate this VTQA dataset, comparing the performance of both state-of-the-art VQA models and our proposed base-line model, the Key Entity Cross-Media Reasoning Network (KECMRN). The VTQA task poses formidable challenges for traditional VQA models, underscoring its intrinsic complexity. Conversely, KECMRN exhibits a modest improvement, signifying its potential in multimedia entity alignment and multi-step reasoning. Our analysis underscores the diversity, difficulty, and scale of the VTQA task compared to previous multimodal QA datasets. In conclusion, we anticipate that this dataset will serve as a pivotal resource for advancing and evaluating models proficient in multime-dia entity alignment, multi-step reasoning, and open-ended answer generation. Our dataset and code is available at https://visual-text-qa.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following BehaviorWeikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li 等CVPR 2026 · 被引用 6 次
- From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic ReasoningYuhui Zeng, Haoxiang Wu, Wenjie Nie, Guangyao Chen 等ICCV 2025 · 被引用 2 次
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music ReasoningGagan Mundada, Yash Vishe, Amit Namburi, Xin Xu 等EMNLP 2025 · 被引用 1 次
- RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and ManipulationLinfei Li, Lin Zhang, Ying ShenCVPR 2026
- Attention Grounded Enhancement for Visual Document RetrievalWanqing Cui, Wei Huang, Yazhi Guo, Yibo Hu 等SIGIR 2026
它引用的顶会 Paper1
相关 Paper
- CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop ReasoningJunyoung Sung, Seungwoo Lyu, Minjun Kim, Sumin An 等CVPR 2026 · 被引用 2 次
- ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringDuong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le PhuocICCV 2025 · 被引用 2 次
- MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingRevanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin 等AAAI 2022 · 被引用 37 次
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao 等ACL 2026
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 被引用 19 次
