Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered Pipelining
Zili Zhang, Fangyue Liu, Gang Huang, Xuanzhe Liu, Xin Jin
Abstract
Vector query processing powers a wide range of AI applications. While GPUs are optimized for massive vector operations, today's practice relies on CPUs to process queries for large vector datasets, due to limited GPU memory.
We present RUMMY, the first GPU-accelerated vector query processing system that achieves high performance and supports large vector datasets beyond GPU memory. The core of RUMMY is a novel reordered pipelining technique that exploits the characteristics of vector query processing to efficiently pipeline data transmission from host memory to GPU memory and query processing in GPU. Specifically, it leverages three ideas: (i) cluster-based retrofitting to eliminate redundant data transmission across queries in a batch, (ii) dynamic kernel padding with cluster balancing to maximize spatial and temporal GPU utilization for GPU computation, and (iii) query-aware reordering and grouping to optimally overlap transmission and computation. We also tailor GPU memory management for vector queries to reduce GPU memory fragmentation and cache misses. We evaluate RUMMY with a variety of billion-scale benchmarking datasets. The experimental results show that RUMMY outperforms IVF-GPU with CUDA unified memory by up to 135×. Compared to the CPU-based solution (with 64 vCPUs), RUMMY (with one NVIDIA A100 GPU) achieves up to 23.1× better performance and is up to 37.7× more cost-effective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 083bf1a0-e770-4f49-9be9-b5b7a6df411fCited by top-tier papers18
- Towards High-throughput and Low-latency Billion-scale Vector Search via CPU/GPU Collaborative Filtering and Re-rankingBing Tian, Haikun Liu, Yuhang Tang, Shihai Xiao et al.FAST 2025 · 49 citations
- Turbocharging Vector Databases using Modern SSDsJoobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do et al.VLDB 2025 · 13 citations
- ANSMET: Approximate Nearest Neighbor Search with Near-Memory Processing and Hybrid Early TerminationYiwei Li, Yuxin Jin, Boyu Tian, Huanchen Zhang et al.ISCA 2025 · 10 citations
- Turbocharge ANNS on Real Processing-in-Memory by Enabling Fine-Grained Per-PIM-Core SchedulingPuqing Wu, Minhui Xie, Enrui Zhao, Dafang Zhang et al.USENIX ATC 2025 · 8 citations
- CoTra: Towards Efficient and Scalable Distributed Vector Search with RDMAXiangyu Zhi, Meng Chen, Xiao Yan, Baotong Lu et al.SIGMOD 2026 · 7 citations
Builds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 152 citations
- SONG: Approximate Nearest Neighbor Search on GPUWeijie Zhao, Shulong Tan, Ping LiICDE 2020 · 103 citations
Related papers
- Hitcher: Efficient GPU-based Vector Search via Cluster-Centric Kernel and Hitch-Ride OrderingQihui Zhou, Changji Li, Guanxian Jiang, Chenhao Ma et al.KDD 2026
- GaccO - A GPU-accelerated OLTP DBMSNils Boeschen, Carsten BinnigSIGMOD 2022 · 18 citations
- VecFlow: A High-Performance Vector Data Management System for Filtered-Search on GPUsJingyi Xi, Chenghao Mo, Ben Karsin, Artem M. Chirkin et al.SIGMOD 2026 · 1 citation
- SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector SearchYuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun et al.VLDB 2026 · 1 citation
- GPU-Native Approximate Nearest Neighbor Search with IVF-RaBitQ: Fast Index Build and SearchJifan Shi, Jianyang Gao, James Xia, Tamas B. Fehér et al.VLDB 2026 · 7 citations
