CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQA
Jinxu Zhang, Yongqi Yu, Yu Zhang
Abstract
Document Visual Question Answering (DVQA) involves responding to queries based on the contents of document images. Existing works are confined to locating information within a single page and lack support for cross-page question-and-answer interactions. Furthermore, the token length limitation on model inputs can lead to the truncation of answer-relevant segments. In this study, we present CREAM, an innovative methodology that focuses on high-performance retrieval and integrates relevant multimodal document information to effectively address this critical issue. To overcome the limitations of current text embedding similarity methods, we first employ a coarse-to-fine retrieval and ranking approach. The coarse phase calculates the similarity between the query and text chunk embeddings, while the fine phase involves multiple rounds of grouping and ordering with a large language model to identify the text chunks most relevant to the query. Subsequently, integrating an attention pooling mechanism for multi-page document images into the vision encoder allows us to effectively merge the visual information of multi-page documents, enabling the multimodal large language model (MLLM) to simultaneously process both single-page and multi-page documents. Finally, we apply various parameter-efficient tuning methods to enhance document visual question-answering performance. Experiments demonstrate that our approach secures state-of-the-art results across various document datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang et al.CVPR 2026 · 9 citations
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang et al.ACL 2026 · 2 citations
- URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingYongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng et al.AAAI 2026 · 2 citations
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingChao Deng, Jiale Yuan, Pi Bu, Peijie Wang et al.ACL 2025
Related papers
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- Attention as Selector: Unlocking VLM Attention for Long Document Page RetrievalMinfeng Zhu, Linxin Bao, Wei Chen, Linchao ZhuACL 2026
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu et al.EMNLP 2025 · 1 citation
- GRAM: Global Reasoning for Multi-Page VQATsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts et al.CVPR 2024
- Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsGeewook Kim, Hodong Lee, Daehee Kim, Haeji Jung et al.EMNLP 2023 · 2 citations
