SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, Kuniko Saito
摘要
Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets have been proposed for developing document VQA systems, most of the existing datasets focus on understanding the content relationships within a single image and not across multiple images. In this study, we propose a new multi-image document VQA dataset, SlideVQA, containing 2.6k+ slide decks composed of 52k+ slide images and 14.5k questions about a slide deck. SlideVQA requires complex reasoning, including single-hop, multi-hop, and numerical reasoning, and also provides annotated arithmetic expressions of numerical answers for enhancing the ability of numerical reasoning. Moreover, we developed a new end-to-end document VQA model that treats evidence selection and question answering as a unified sequence-to-sequence format. Experiments on SlideVQA show that our model outperformed existing state-of-the-art QA models, but that it still has a large gap behind human performance. We believe that our dataset will facilitate research on document VQA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao 等ICLR 2024 · 被引用 95 次
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen 等NeurIPS 2025 · 被引用 76 次
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang 等NeurIPS 2025 · 被引用 69 次
- InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with InstructionsRyota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito 等AAAI 2024 · 被引用 39 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav 等ICLR 2021 · 被引用 229 次
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 被引用 201 次
相关 Paper
- Towards Complex Document Understanding By Discrete ReasoningFengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang 等ACM MM 2022 · 被引用 40 次
- ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringDuong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le PhuocICCV 2025 · 被引用 2 次
- DOC2PPT: Automatic Presentation Slides Generation from Scientific DocumentsTsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale SongAAAI 2022 · 被引用 83 次
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningKang Chen, Xiangqian WuCVPR 2024
- Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented GenerationPeiyang Liu, Ziqiang Cui, Xi Wang, Di Liang 等SIGIR 2026
