DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQA
Jinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu Zhang
Abstract
Understanding the content of multi-page documents with rich layout information is a challenging task. Recent multimodal large language models (MLLMs) have made remarkable progress in understanding single-page document images. However, the understanding of multi-page documents remains insufficiently explored. This work proposes a Document Retrieval-enhanced, Expert-guided, Attention-aware Multimodal Framework, dubbed DREAM. Specifically, we propose a confidence-based, high-level semantic, multimodal retrieval method. Then, we propose a machine learning algorithm to complement the result of confidence-based retrieval and multimodal embedding similarity retrieval to obtain the most query-relevant set of document images. Subsequently, we designed a decoupled cross-page attention-aware multimodal language model for multi-page documents to interpret these retrieved images and produce the final answer. Experimental results demonstrate the effectiveness of the retrieval module within the framework, as well as the robust performance of the multimodal model in multi-page document comprehension. These findings offer a compelling solution for multi-page document comprehension and cross-page document visual question answering.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5720d4ec-684d-4986-b05f-2428236dd640Related papers
- CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQAJinxu Zhang, Yongqi Yu, Yu ZhangACM MM 2024 · 4 citations
- Attention as Selector: Unlocking VLM Attention for Long Document Page RetrievalMinfeng Zhu, Linxin Bao, Wei Chen, Linchao ZhuACL 2026
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu et al.EMNLP 2025 · 1 citation
- SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document UnderstandingJian Chen, Ruiyi Zhang, Yufan Zhou, Tong Yu et al.ICLR 2025
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document UnderstandingKeliang Liu, Zizhi Chen, Mingcheng Li, Jingqun Tang et al.CVPR 2026 · 19 citations
