Neos: A NVMe-GPUs Direct Vector Service Buffer in User Space
Yuchen Huang, Xiaopeng Fan, Song Yan, Chuliang Weng
摘要
With the development of AI generated content and LLM (Large Language Model), demands of vector management have brought prosperity to vector databases. However, the status that vectors cannot be retrieved before being indexed, harms timeliness of vector databases. Updating indexes immediately when adding new vectors, reduces throughput of storage. Due to this contradiction, when facing streaming data, using vector database solely in vector services cannot have it both ways: real-time searches and high-throughput storage. This paper proposes a vector buffer engine, Neos. It is designed for real-time unindexed-vector searches on streaming input and buffering vectors with high throughput before loading them into vector databases. On one hand, we build a lightweight storage on raw NVMe device and liberate throughput from indexes, to maximize storage performance. On the other hand, we realize direct NVMe-GPUs 110 stack and a CPU-GPU heterogeneous task architecture for low-latency unindexed-vector searches on streaming data. Experiments show that our approach performs with 1.5x to 3.4x bandwidth, as low as 20% latency compared to existing 110 stacks, and up to orders-of-magnitude higher vector storage throughput under concurrent RIW workloads. Further, N eos can handle real-time unindexed - vector searches with millisecond-level latency on streaming input, a capability that current vector systems lack.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Turbocharging Vector Databases using Modern SSDsJoobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do 等VLDB 2025 · 被引用 13 次
- Accelerating Graph Indexing for ANNS on Modern CPUsMengzhao Wang, Haotian Wu, Xiangyu Ke, Yunjun Gao 等SIGMOD 2025 · 被引用 6 次
- Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor SearchZhonggen Li, Xiangyu Ke, Yifan Zhu, Bocheng Yu 等SIGMOD 2026 · 被引用 5 次
- Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPUChangyue Liao, Mo Sun, Zihan Yang, Jun Xie 等ICDE 2025 · 被引用 4 次
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun 等ICDE 2025 · 被引用 4 次
相关 Paper
- VStream: A Distributed Streaming Vector Search SystemShenghao Gong, Haobo Sun, Ziquan Fang, Liu Liu 等VLDB 2025 · 被引用 4 次
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAGJunkyum Kim, Divya MahajanHPCA 2026 · 被引用 2 次
- SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector SearchYuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun 等VLDB 2026 · 被引用 1 次
- NeuVSA: A Unified and Efficient Accelerator for Neural Vector SearchZiming Yuan, Lei Dai, Wen Li, Jie Zhang 等HPCA 2025 · 被引用 1 次
- VStore: in-storage graph based vector search acceleratorShengwen Liang, Ying Wang, Ziming Yuan, Cheng Liu 等DAC 2022 · 被引用 20 次
