Lune

ACM MM2025顶会

DREAM: Integrating Hierarchical Multimodal Retrieval with Multi-page Multimodal Language Model for Documents VQA

Jinxu Zhang, Qiyuan Fan, Yongqi Yu, Yu Zhang

2025年份

摘要

Understanding the content of multi-page documents with rich layout information is a challenging task. Recent multimodal large language models (MLLMs) have made remarkable progress in understanding single-page document images. However, the understanding of multi-page documents remains insufficiently explored. This work proposes a Document Retrieval-enhanced, Expert-guided, Attention-aware Multimodal Framework, dubbed DREAM. Specifically, we propose a confidence-based, high-level semantic, multimodal retrieval method. Then, we propose a machine learning algorithm to complement the result of confidence-based retrieval and multimodal embedding similarity retrieval to obtain the most query-relevant set of document images. Subsequently, we designed a decoupled cross-page attention-aware multimodal language model for multi-page documents to interpret these retrieved images and produce the final answer. Experimental results demonstrate the effectiveness of the retrieval module within the framework, as well as the robust performance of the multimodal model in multi-page document comprehension. These findings offer a compelling solution for multi-page document comprehension and cross-page document visual question answering.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖