MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
Zilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen, Xintao Chen, Vicente Ordonez, Vijai Mohan
Abstract
Universal multimodal embedding models have achieved great success in capturing semantic relevance between queries and candidates. However, current methods either condense queries and candidates into a single vector, potentially limiting the expressiveness for fine-grained information, or produce too many vectors that are prohibitively expensive for multi-vector retrieval. In this work, we introduce MetaEmbed, a new framework for multimodal retrieval that rethinks how multimodal embeddings are constructed and interacted with at scale. During training, a fixed number of learnable Meta Tokens are appended to the input sequence. At test-time, their last-layer contextualized representations serve as compact yet expressive multi-vector embeddings. Through the proposed Matryoshka Multi-Vector Retrieval training, MetaEmbed learns to organize information by granularity across multiple vectors. As a result, we enable test-time scaling in multimodal retrieval where users can balance retrieval quality against efficiency demands by selecting the number of tokens used for indexing and retrieval interactions. Extensive evaluations on the Massive Multimodal Embedding Benchmark (MMEB) and the Visual Document Retrieval Benchmark (ViDoRe) confirm that MetaEmbed achieves state-of-the-art retrieval performance while scaling robustly to models with 32B parameters. Code is available at https://github.com/facebookresearch/MetaEmbed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 763b75bb-ff52-4091-8bb9-72f1a7d89925Cited by top-tier papers5
- ModernVBERT: Towards Smaller Visual Document RetrieversPaul Teiletche, Quentin Macé, Max Conti, António Loison et al.ICML 2026 · 17 citations
- Multi-Vector Index Compression in Any ModalityHanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo et al.SIGIR 2026 · 11 citations
- ReMatch: Boosting Representation through Matching for Multimodal RetrievalQianying Liu, Xiao Liang, Zhiqiang Zhang, Yibo Chen et al.CVPR 2026 · 8 citations
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan et al.ACL 2026 · 7 citations
- Efficient and High-Fidelity Omni Modality RetrievalChuong Huynh, Manh Luong, Abhinav ShrivastavaCVPR 2026 · 3 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding TasksZiyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz et al.ICLR 2025
- Towards Text-Image Interleaved RetrievalXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang et al.ACL 2025 · 1 citation
- Mieb: Massive Image Embedding BenchmarkChenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling et al.ICCV 2025
- Matryoshka Multimodal ModelsMu Cai, Jianwei Yang, Jianfeng Gao, Yong Jae LeeICLR 2025
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
