mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou
2025年份
32顶会引用
摘要
image Continue-pretraining, and Multi-task Finetuning. DocOwl2 sets a new state-of-theart across multi-page document understanding benchmarks and reduces first token latency by more than 50%. Compared to single-image MLLMs trained on similar data, our DocOwl2 achieves comparable single-page understanding performance with less than 20% of the visual tokens. Our codes, models, and data are released at https://github.com/X-PLUG/ mPLUG-DocOwl/tree/main/DocOwl2 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu 等NeurIPS 2025 · 被引用 32 次
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison 等ICML 2026 · 被引用 17 次
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang 等CVPR 2026 · 被引用 9 次
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla 等ICLR 2026 · 被引用 9 次
- A Token-Level Text Image Foundation Model for Document UnderstandingTongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
相关 Paper
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham 等CVPR 2025
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingWenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang 等CVPR 2025
- Should VLMs be Pre-trained with Image Data?Sedrick Keh, Jean Mercat, Samir Yitzhak Gadre, Kushal Arora 等ICLR 2025
- DocR1: Evidence Page-Guided GRPO for Multi-Page Document UnderstandingJunyu Xiong, Yonghui Wang, Weichao Zhao, Chenyu Liu 等AAAI 2026 · 被引用 5 次
- Hierarchical Visual Feature Aggregation for OCR-Free Document UnderstandingJaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung HanNeurIPS 2024 · 被引用 19 次
