Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
Yuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang, Yuyi Zhang, Wenyu Ruan, Xiaojin Zhang, Zhongyu Wei, Zhenbo Luo, Jian Luan, Wei Chen, Xiang Bai
Abstract
Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existing OCR-free methods face a trade-off between capacity and precision: end-to-end models scale poorly with document length, while visual retrieval-based pipelines are brittle and passive. We propose Doc-, an OCR-free agentic framework that casts multi-page DocVQA as sequential evidence aggregation. Doc- begins with a thumbnail overview, then actively navigates via semantic retrieval and targeted page fetching, and aggregates evidence in a structured working memory for grounded reasoning. Trained by imitation learning from expert trajectories and further optimized with Group Relative Policy Optimization, Doc- balances answer accuracy with evidence-seeking efficiency. Across five benchmarks, Doc- outperforms open-source baselines and approaches proprietary models, improving out-of-domain performance by up to 47.9% over RAG baseline. Other results reveal effective evidence aggregation with selective attention, not increased input pages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3d51886-2382-429d-b11e-5bd0c76e7a2cCited by top-tier papers1
Ask how each one uses itBuilds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa et al.AAAI 2023 · 178 citations
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement LearningQiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen et al.NeurIPS 2025 · 76 citations
Related papers
- UniDocVLM: Enhancing Visual Reasoning and Document Understanding for VLM via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Lang Chen et al.KDD 2026
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu et al.EMNLP 2025 · 1 citation
- ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question AnsweringAymen Lassoued, Mohamed Ali Souibgui, Yousri KessentiniCVPR 2026 · 1 citation
- LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document UnderstandingZhivar Sourati, Zheng Wang, Marianne Menglin Liu, Yazhe Hu et al.ACL 2026 · 5 citations
- GRAM: Global Reasoning for Multi-Page VQATsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts et al.CVPR 2024
