mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou
2025Year
32Top-tier citations
Abstract
image Continue-pretraining, and Multi-task Finetuning. DocOwl2 sets a new state-of-theart across multi-page document understanding benchmarks and reduces first token latency by more than 50%. Compared to single-image MLLMs trained on similar data, our DocOwl2 achieves comparable single-page understanding performance with less than 20% of the visual tokens. Our codes, models, and data are released at https://github.com/X-PLUG/ mPLUG-DocOwl/tree/main/DocOwl2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu et al.NeurIPS 2025 · 32 citations
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison et al.ICML 2026 · 17 citations
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang et al.CVPR 2026 · 9 citations
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla et al.ICLR 2026 · 9 citations
- A Token-Level Text Image Foundation Model for Document UnderstandingTongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo et al.ICCV 2025 · 5 citations
Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
Related papers
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham et al.CVPR 2025
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingWenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang et al.CVPR 2025
- Should VLMs be Pre-trained with Image Data?Sedrick Keh, Jean Mercat, Samir Yitzhak Gadre, Kushal Arora et al.ICLR 2025
- DocR1: Evidence Page-Guided GRPO for Multi-Page Document UnderstandingJunyu Xiong, Yonghui Wang, Weichao Zhao, Chenyu Liu et al.AAAI 2026 · 5 citations
- Hierarchical Visual Feature Aggregation for OCR-Free Document UnderstandingJaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung HanNeurIPS 2024 · 19 citations
