MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval
Junjie Zhou, Yongping Xiong, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian
Abstract
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70× more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, welltrained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 893c2c71-1770-417a-9598-92b83ca91a0eCited by top-tier papers34
- Think Then Embed: Generative Context Improves Multimodal EmbeddingXuanming Cui, Jianpeng Cheng, Hong-You Chen, Satya Narayan Shukla et al.ICLR 2026 · 41 citations
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen et al.ICLR 2026 · 40 citations
- UME-R1: Exploring Reasoning-Driven Generative Multimodal EmbeddingsZhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou et al.ICLR 2026 · 38 citations
- Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningRaghuveer Thirukovalluru, Rui Meng, Ye Liu, Karthikeyan K et al.NeurIPS 2025 · 30 citations
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsHaonan Chen, Hong Liu, Yuping Luo, Liang Wang et al.ACL 2026 · 20 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah et al.CVPR 2025
- Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and NegativesZhangchi Feng, Richong Zhang, Zhijie NieACM MM 2024 · 14 citations
- CoVR: Learning Composed Video Retrieval from Web Video CaptionsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolAAAI 2024 · 81 citations
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 120 citations
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
