VecFlow-Chamfer: A GPU-based Data Management System for High-Performance Multi-Vector Search on Superchips
Chenghao Mo, Ben Karsin, Philip Adams, Minjia Zhang
Abstract
Many emerging AI applications demand retrieval systems that go beyond document-level relevance and capture token-level semantics. Multi-vector search addresses this need through fine-grained semantic matching between token-level embeddings of queries and documents. However, it introduces significant system challenges, including compute-intensive set-to-set scoring, complex candidate filtering, and high memory overhead from storing dense token-level embeddings. Prior systems mitigate these challenges through GPU-based similarity calculation and indexing structures (e.g., IVFPQ-GPU) but often at the cost of reduced retrieval accuracy and suffer from low GPU utilization. We present , a GPU-based vector data management system that enables low-latency and high-recall multi-vector search on modern Superchip architectures. achieves this through a combination of three novel optimizations: (1) , a GPU-native, compression-free index tailored for multi-vector search with fine-grained anchor vectors enabling scalable index construction and low-latency, high-accuracy candidate generation via CAGRA-based GPU routing; (2) , a highly-optimized GPU kernel that enables single-digit millisecond Chamfer scoring over tens of thousands of candidate documents; and (3) , a tiered vector storage layer that supports on-demand, low-latency access to full-precision embeddings across Grace-Hopper NVLink-C2C interconnects. Together, these techniques enable to perform multi-vector search over hundreds of millions of document token embeddings with unprecedented high recall and low latency. Compared to state-of-the-art systems like PLAID and MUVERA, achieves an order-of-magnitude lower latency while significantly improving recall.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0a340c65-ff8d-4bc8-87c2-7d54f1021f57Related papers
- GIGP+: A CPU-GPU Co-Processing Engine for Multi-Vector RetrievalZheng Bian, Man Lung Yiu, Bo TangSIGIR 2026
- MUVERA: Multi-Vector Retrieval via Fixed Dimensional EncodingLaxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee et al.NeurIPS 2024 · 56 citations
- SIVF: GPU-Resident IVF Index for Streaming Vector AnalyticsDongfang ZhaoHPDC 2026
- VStore: in-storage graph based vector search acceleratorShengwen Liang, Ying Wang, Ziming Yuan, Cheng Liu et al.DAC 2022 · 20 citations
- Hitcher: Efficient GPU-based Vector Search via Cluster-Centric Kernel and Hitch-Ride OrderingQihui Zhou, Changji Li, Guanxian Jiang, Chenhao Ma et al.KDD 2026
