MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encoding
Laxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee, Vahab Mirrokni
摘要
Neural embedding models have become a fundamental component of modern information retrieval (IR) pipelines. These models produce a single embedding per data-point, allowing for fast retrieval via highly optimized maximum inner product search (MIPS) algorithms. Recently, beginning with the landmark ColBERT paper, multi-vector models, which produce a set of embedding per data point, have achieved markedly superior performance for IR tasks. Unfortunately, using these models for IR is computationally expensive due to the increased complexity of multi-vector retrieval and scoring. In this paper, we introduce MUVERA (MUlti-VEctor Retrieval Algorithm), a retrieval mechanism which reduces multi-vector similarity search to single-vector similarity search. This enables the usage of off-the-shelf MIPS solvers for multi-vector retrieval. MUVERA asymmetrically generates Fixed Dimensional Encodings (FDEs) of queries and documents, which are vectors whose inner product approximates multi-vector similarity. We prove that FDEs give high-quality -approximations, thus providing the first single-vector proxy for multi-vector similarity with theoretical guarantees. Empirically, we find that FDEs achieve the same recall as prior state-of-the-art heuristics while retrieving 2-5 fewer candidates. Compared to prior state of the art implementations, MUVERA achieves consistently good end-to-end recall and latency across a diverse set of the BEIR retrieval datasets, achieving an average of 10 improved recall with lower latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen 等ICLR 2026 · 被引用 40 次
- IGP: Efficient Multi-Vector Retrieval via Proximity Graph IndexZheng Bian, Man Lung Yiu, Bo TangSIGIR 2025 · 被引用 5 次
- LEMUR: Learned Multi-Vector RetrievalElias Jääsaari, Ville Hyvönen, Teemu RoosICML 2026 · 被引用 3 次
- Distance Adaptive Beam Search for Provably Accurate Graph-Based Nearest Neighbor SearchYousef Al-Jazzazi, Haya Diwan, Jinrui Gou, Cameron Musco 等NeurIPS 2025 · 被引用 3 次
- GEM: A Native Graph-based Index for Multi-Vector RetrievalYao Tian, Zhoujin Tian, Xi Zhao, Ruiyuan Zhang 等SIGMOD 2026 · 被引用 2 次
它引用的顶会 Paper14
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- Baleen: Robust Multi-Hop Reasoning at Scale via Condensed RetrievalOmar Khattab, Christopher Potts, Matei A. ZahariaNeurIPS 2021 · 被引用 92 次
- Rethinking the Role of Token Retrieval in Multi-Vector RetrievalJinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 等NeurIPS 2023 · 被引用 60 次
相关 Paper
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector RetrievalLixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng 等ICML 2026
- VecFlow-Chamfer: A GPU-based Data Management System for High-Performance Multi-Vector Search on SuperchipsChenghao Mo, Ben Karsin, Philip Adams, Minjia ZhangSIGMOD 2026 · 被引用 2 次
- GIGP+: A CPU-GPU Co-Processing Engine for Multi-Vector RetrievalZheng Bian, Man Lung Yiu, Bo TangSIGIR 2026
- Pseudo-Relevance for Enhancing Document RepresentationJihyuk Kim, Seung-won Hwang, Seoho Song, Hyeseon Ko 等EMNLP 2022 · 被引用 1 次
- WARP: An Efficient Engine for Multi-Vector RetrievalJan Luca Scheerer, Matei Zaharia, Christopher Potts, Gustavo Alonso 等SIGIR 2025 · 被引用 8 次
