OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu
摘要
Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the narrow coverage of document types and the simplified, unrealistic evaluation procedures in existing benchmarks. To address these gaps, we introduce OmniDocBench, a novel benchmark featuring high-quality annotations across nine document sources, including academic papers, textbooks, and more challenging cases such as handwritten notes and densely typeset newspapers. OmniDocBench supports flexible, multi-level evaluations-ranging from an end-to-end assessment to the task-specific and attribute-based analysis-using 19 layout categories and 15 attribute labels. We conduct a thorough evaluation of both pipeline-based methods and endto-end vision-language models, revealing their strengths and weaknesses across different document types. Om-niDocBench sets a new standard for the fair, diverse, and fine-grained evaluation in document parsing. Dataset and code are available at https://github.com/ opendatalab/OmniDocBench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR TasksCheng Cui, yubo zhang, Ting Sun, Xueqing Wang 等CVPR 2026 · 被引用 11 次
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationJunyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang 等ICCV 2025 · 被引用 10 次
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware TrainingGengluo Li, Pengyuan Lyu, Chengquan Zhang, Huawen Shen 等CVPR 2026 · 被引用 9 次
- Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCRYufeng Zhong, Lei Chen, Zhixiong Zeng, Xuanle Zhao 等CVPR 2026 · 被引用 8 次
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table RecognitionJunyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 被引用 243 次
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui 等ACM MM 2022 · 被引用 184 次
- SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text RecognitionMingxin Huang, Yuliang Liu, Zhenghao Peng, Chongyu Liu 等CVPR 2022 · 被引用 150 次
相关 Paper
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen 等AAAI 2026
- MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document UnderstandingKetong Chen, Yuhao Chen, Yang XueAAAI 2026 · 被引用 1 次
- OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial DomainShuting Wang, Jiejun Tan, Zhicheng Dou, Ji-Rong WenEMNLP 2025 · 被引用 6 次
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui 等ICLR 2025
- MORE: A Multilingual Document Parsing Benchmark and EvaluationLong Xu, Binghong Wu, TingHao YU, Hao Feng 等ICML 2026 · 被引用 3 次
