UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement Learning
Zongsheng Cao, Anran Liu, Jun Xie, Lang Chen, Feng Chen, Zigan Wang
摘要
Document question answering over scanned pages requires two coupled abilities: (i) canonicalizing complex layouts into a faithful textual structure, and (ii) selecting and reasoning over query-relevant evidence from that structure. Most existing pipelines decouple OCR from retrieval-augmented reasoning and optimize OCR for global reconstruction, which often misaligns with evidence needs and causes brittle grounding in multi-page settings. We propose UniDocVLM, an end-to-end framework that unifies OCR and visual RAG within a single vision-language model: the model first generates a structured parse of retrieved pages, then activates question-relevant evidence from the parse to support grounded reasoning and answering. To train UniDocVLM under heterogeneous supervision, we introduce a unified JR-GRPO reinforcement learning recipe with lightweight, verifiable rewards, including format, layout-aware OCR, evidence-consistency, and answer-correctness signals, and route them to the corresponding parts of the output to improve credit assignment and reduce interference. Experiments on multi-page document QA benchmarks show that UniDocVLM yields more reliable evidence grounding and improves downstream accuracy under complex layouts.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang 等ACL 2026 · 被引用 2 次
- DocR1: Evidence Page-Guided GRPO for Multi-Page Document UnderstandingJunyu Xiong, Yonghui Wang, Weichao Zhao, Chenyu Liu 等AAAI 2026 · 被引用 5 次
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu 等EMNLP 2025 · 被引用 1 次
- ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained AlignmentHao Yang, Yifan Ji, Zhipeng Xu, Zhenghao Liu 等SIGIR 2026 · 被引用 4 次
- CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic FrameworkYuexi Du, Jinglu Wang, Shujie Liu, Nicha C. Dvornek 等ICLR 2026 · 被引用 4 次
