Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models
Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang
摘要
Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt multimodal large language models (MLLMs) to realize UMR using only text data. However, our preliminary experiments demonstrate that more diverse multimodal training data can further unlock the potential of MLLMs. Despite its effectiveness, the existing multimodal training data is highly imbalanced in terms of modality, which motivates us to develop a training data synthesis pipeline and construct a large-scale, highquality fused-modal training dataset. Based on the synthetic training data, we develop the General Multimodal Embedder (GME), an MLLM-based dense retriever designed for UMR. Furthermore, we construct a comprehensive UMR Benchmark (UMRB) to evaluate the effectiveness of our approach. Experimental results show that our method achieves state-of-the-art performance among existing UMR methods. Last, we provide in-depth analyses of model scaling and training strategies, and perform ablation studies on both the model and synthetic data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- UME-R1: Exploring Reasoning-Driven Generative Multimodal EmbeddingsZhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou 等ICLR 2026 · 被引用 38 次
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and GranularitiesWoongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek 等ACL 2026 · 被引用 14 次
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product UnderstandingZhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu 等CVPR 2026 · 被引用 9 次
- Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM RerankingXin Zhang, Ziqi Dai, Mingxin Li, Yanzhao Zhang 等ICLR 2026 · 被引用 8 次
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan 等ACL 2026 · 被引用 7 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
相关 Paper
- Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMSSheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin 等ICLR 2025
- Towards Text-Image Interleaved RetrievalXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang 等ACL 2025 · 被引用 1 次
- TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal ModelsLeigang Qu, Haochuan Li, Tan Wang, Wenjie Wang 等ICLR 2025
- MegaPairs: Massive Data Synthesis for Universal Multimodal RetrievalJunjie Zhou, Yongping Xiong, Zheng Liu, Ze Liu 等ACL 2025
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMsXiaojie Li, Chu Li, Shi-Zhe Chen, Xi ChenICLR 2026 · 被引用 10 次
