MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
Ketong Chen, Yuhao Chen, Yang Xue
摘要
Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature simplistic layouts, and support limited tasks. Consequently, they fail to evaluate model performance for Visually Rich Document Understanding (VRDU), a critical challenge involving complex layouts and dense text. To address this, we introduce DocWeaver, a novel multi-agent pipeline that leverages Large Language Models to automatically generate a new benchmark. The result is MosaicDoc, a large-scale, bilingual (Chinese and English) resource designed to push the boundaries of VRDU. Sourced from newspapers and magazines, MosaicDoc features diverse and complex layouts (including multi-column and non-Manhattan), rich stylistic variety from 196 publishers, and comprehensive multi-task annotations (OCR, VQA, reading order, and localization). With 72K images and over 600K QA pairs, MosaicDoc serves as a definitive benchmark for the field. Our extensive evaluation of state-of-the-art models on this benchmark reveals their current limitations in handling real-world document complexity and charts a clear path for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 被引用 201 次
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang 等AAAI 2025 · 被引用 138 次
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- LayoutReader: Pre-training of Text and Layout for Reading Order DetectionZilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang 等EMNLP 2021 · 被引用 51 次
- M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisHiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang 等CVPR 2023
相关 Paper
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu 等CVPR 2025
- MORE: A Multilingual Document Parsing Benchmark and EvaluationLong Xu, Binghong Wu, TingHao YU, Hao Feng 等ICML 2026 · 被引用 3 次
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingChao Deng, Jiale Yuan, Pi Bu, Peijie Wang 等ACL 2025
- FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step ComputationZichen Tang, Haihong E, Rongjin Li, Jiacheng Liu 等AAAI 2026
- Cross-Lingual Text-Rich Visual Comprehension: An Information Theory PerspectiveXinmiao Yu, Xiaocheng Feng, Yun Li, Minghui Liao 等AAAI 2025 · 被引用 7 次
