CREAM: Coarse-to-Fine Retrieval and Multi-modal Efficient Tuning for Document VQA
Jinxu Zhang, Yongqi Yu, Yu Zhang
摘要
Document Visual Question Answering (DVQA) involves responding to queries based on the contents of document images. Existing works are confined to locating information within a single page and lack support for cross-page question-and-answer interactions. Furthermore, the token length limitation on model inputs can lead to the truncation of answer-relevant segments. In this study, we present CREAM, an innovative methodology that focuses on high-performance retrieval and integrates relevant multimodal document information to effectively address this critical issue. To overcome the limitations of current text embedding similarity methods, we first employ a coarse-to-fine retrieval and ranking approach. The coarse phase calculates the similarity between the query and text chunk embeddings, while the fine phase involves multiple rounds of grouping and ordering with a large language model to identify the text chunks most relevant to the query. Subsequently, integrating an attention pooling mechanism for multi-page document images into the vision encoder allows us to effectively merge the visual information of multi-page documents, enabling the multimodal large language model (MLLM) to simultaneously process both single-page and multi-page documents. Finally, we apply various parameter-efficient tuning methods to enhance document visual question-answering performance. Experiments demonstrate that our approach secures state-of-the-art results across various document datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document UnderstandingHao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang 等CVPR 2026 · 被引用 9 次
- Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQAYuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang 等ACL 2026 · 被引用 2 次
- URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingYongxin Shi, Jiapeng Wang, Zeyu Shan, Dezhi Peng 等AAAI 2026 · 被引用 2 次
- LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and LocatingChao Deng, Jiale Yuan, Pi Bu, Peijie Wang 等ACL 2025
相关 Paper
- DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQAJinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu ZhangACM MM 2025
- Attention as Selector: Unlocking VLM Attention for Long Document Page RetrievalMinfeng Zhu, Linxin Bao, Wei Chen, Linchao ZhuACL 2026
- SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative RefinementChelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu 等EMNLP 2025 · 被引用 1 次
- GRAM: Global Reasoning for Multi-Page VQATsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts 等CVPR 2024
- Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language ModelsGeewook Kim, Hodong Lee, Daehee Kim, Haeji Jung 等EMNLP 2023 · 被引用 2 次
