Unifying Multimodal Retrieval via Document Screenshot Embedding
Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, Jimmy Lin
摘要
In the real world, documents are organized in different formats and varied modalities. Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for indexing. This process is tedious, prone to errors, and has information loss. To this end, we propose Document Screenshot Embedding (DSE), a novel retrieval paradigm that regards document screenshots as a unified input format, which does not require any content extraction preprocess and preserves all the information in a document (e.g., text, image and layout). DSE leverages a large vision-language model to directly encode document screenshots into dense representations for retrieval. To evaluate our method, we first craft the dataset of Wiki-SS, a 1.3M Wikipedia web page screenshots as the corpus to answer the questions from the Natural Questions dataset. In such a text-intensive document retrieval setting, DSE shows competitive effectiveness compared to other text retrieval methods relying on parsing. For example, DSE outperforms BM25 by 17 points in top-1 retrieval accuracy. Additionally, in a mixed-modality task of slide retrieval, DSE significantly outperforms OCR text retrieval methods by over 15 points in nDCG@10. These experiments show that DSE is an effective document retrieval paradigm for diverse types of documents. Model checkpoints, code, and Wiki-SS collection are released at http://tevatron.ai .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- On the Theoretical Limitations of Embedding-Based RetrievalOrion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk LeeICLR 2026 · 被引用 138 次
- VISA: Retrieval Augmented Generation with Visual Source AttributionXueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon 等ACL 2025 · 被引用 24 次
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsHaonan Chen, Hong Liu, Yuping Luo, Liang Wang 等ACL 2026 · 被引用 20 次
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison 等ICML 2026 · 被引用 17 次
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal RetrievalSiyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao 等ICLR 2026 · 被引用 13 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information RetrievalZheng Liu, Ze Liu, Zhengyang Liang, Junjie Zhou 等ACL 2025 · 被引用 9 次
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich DocumentsRyota Tanaka, Taichi Iki, Taku Hasegawa, Kyosuke Nishida 等CVPR 2025
- Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document RetrievalHao Sun, Yingyan Hou, Jiayan Guo, Bo Wang 等ACL 2025
- Enhancing Vision-Language Pre-Training with Rich SupervisionsYuan Gao, Kunyu Shi, Pengkai Zhu, Edouard Belval 等CVPR 2024
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui 等ICLR 2025
