UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement Learning
Zongsheng Cao, Anran Liu, Jun Xie, Lang Chen, Feng Chen, Zigan Wang
Abstract
Document question answering over scanned pages requires two coupled abilities: (i) canonicalizing complex layouts into a faithful textual structure, and (ii) selecting and reasoning over query-relevant evidence from that structure. Most existing pipelines decouple OCR from retrieval-augmented reasoning and optimize OCR for global reconstruction, which often misaligns with evidence needs and causes brittle grounding in multi-page settings. We propose UniDocVLM, an end-to-end framework that unifies OCR and visual RAG within a single vision-language model: the model first generates a structured parse of retrieved pages, then activates question-relevant evidence from the parse to support grounded reasoning and answering. To train UniDocVLM under heterogeneous supervision, we introduce a unified JR-GRPO reinforcement learning recipe with lightweight, verifiable rewards, including format, layout-aware OCR, evidence-consistency, and answer-correctness signals, and route them to the corresponding parts of the output to improve credit assignment and reduce interference. Experiments on multi-page document QA benchmarks show that UniDocVLM yields more reliable evidence grounding and improves downstream accuracy under complex layouts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eefdd200-2ded-438a-90ea-08a6966f503eCited by top-tier papers1
Ask how each one uses itRelated papers
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang et al.ACL 2026 · 2 citations
- DocR1: Evidence Page-Guided GRPO for Multi-Page Document UnderstandingJunyu Xiong, Yonghui Wang, Weichao Zhao, Chenyu Liu et al.AAAI 2026 · 5 citations
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu et al.EMNLP 2025 · 1 citation
- ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained AlignmentHao Yang, Yifan Ji, Zhipeng Xu, Zhenghao Liu et al.SIGIR 2026 · 4 citations
- CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic FrameworkYuexi Du, Jinglu Wang, Shujie Liu, Nicha C. Dvornek et al.ICLR 2026 · 4 citations
