VisualMRC: Machine Reading Comprehension on Document Images
Ryota Tanaka, Kyosuke Nishida, Sen Yoshida
摘要
Recent studies on machine reading comprehension have focused on text-level understanding but have not yet reached the level of human understanding of the visual layout and content of real-world documents. In this study, we introduce a new visual machine reading comprehension dataset, named VisualMRC, wherein given a question and a document image, a machine reads and comprehends texts in the image to answer the question in natural language. Compared with existing visual question answering (VQA) datasets that contain texts in images, VisualMRC focuses more on developing natural language understanding and generation abilities. It contains 30,000+ pairs of a question and an abstractive answer for 10,000+ document images sourced from multiple domains of webpages. We also introduce a new model that extends existing sequence-to-sequence models, pre-trained with largescale text corpora, to take into account the visual layout and content of documents. Experiments with VisualMRC show that this model outperformed the base sequence-to-sequence models and a state-of-the-art VQA model. However, its performance is still below that of humans on most automatic evaluation metrics. The dataset will facilitate research aimed at connecting vision and language understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa 等AAAI 2023 · 被引用 178 次
- SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsDongyang Liu, Renrui Zhang, Longtian Qiu, Siyuan Huang 等ICML 2024 · 被引用 149 次
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- RikiNet: Reading Wikipedia Pages for Natural Question AnsweringDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan 等ACL 2020 · 被引用 55 次
相关 Paper
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji 等EMNLP 2021 · 被引用 42 次
- InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with InstructionsRyota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito 等AAAI 2024 · 被引用 39 次
- Towards Complex Document Understanding By Discrete ReasoningFengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang 等ACM MM 2022 · 被引用 40 次
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
- V-Doc : Visual questions answers with DocumentsYihao Ding, Zhe Huang, Runlin Wang, Yanhang Zhang 等CVPR 2022 · 被引用 19 次
