MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin
Tianshuo Zhou, Sen Mei, Xinze Li, Zhenghao Liu, Chenyan Xiong, Zhiyuan Liu, Yu Gu, Ge Yu
Abstract
This paper proposes Multi-modAl Retrieval model via Visual modulE pLugin (MARVEL), which learns an embedding space for queries and multi-modal documents to conduct retrieval. MARVEL encodes queries and multimodal documents with a unified encoder model, which helps to alleviate the modality gap between images and texts. Specifically, we enable the image understanding ability of the welltrained dense retriever, T5-ANCE, by incorporating the visual module's encoded image features as its inputs. To facilitate the multi-modal retrieval tasks, we build the ClueWeb22-MM dataset based on the ClueWeb22 dataset, which regards anchor texts as queries, and extracts the related text and image documents from anchor-linked web pages. Our experiments show that MARVEL significantly outperforms the state-of-the-art methods on the multi-modal retrieval dataset WebQA and ClueWeb22-MM. MARVEL provides an opportunity to broaden the advantages of text retrieval to the multimodal scenario. Besides, we also illustrate that the language model has the ability to extract image semantics and partly map the image features to the input word embedding space. All codes are available at https://github. com/OpenMatch/MARVEL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition QueryWei Chow, Yuan Gao, Linfeng Li, Xian Wang et al.NeurIPS 2025 · 6 citations
- Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsZhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang et al.ACM MM 2025 · 4 citations
- Towards Text-Image Interleaved RetrievalXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang et al.ACL 2025 · 1 citation
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
- MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question AnsweringSeokwon Song, Minsu Park, Gunhee KimAAAI 2026
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryZiniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang et al.CVPR 2023
- VISTA: Visualized Text Embedding For Universal Multi-Modal RetrievalJunjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao et al.ACL 2024
- Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal RetrievalZhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu et al.ICLR 2023 · 6 citations
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsYuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao et al.CVPR 2025
- Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalDavide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2025
