Neos: A NVMe-GPUs Direct Vector Service Buffer in User Space
Yuchen Huang, Xiaopeng Fan, Song Yan, Chuliang Weng
Abstract
With the development of AI generated content and LLM (Large Language Model), demands of vector management have brought prosperity to vector databases. However, the status that vectors cannot be retrieved before being indexed, harms timeliness of vector databases. Updating indexes immediately when adding new vectors, reduces throughput of storage. Due to this contradiction, when facing streaming data, using vector database solely in vector services cannot have it both ways: real-time searches and high-throughput storage. This paper proposes a vector buffer engine, Neos. It is designed for real-time unindexed-vector searches on streaming input and buffering vectors with high throughput before loading them into vector databases. On one hand, we build a lightweight storage on raw NVMe device and liberate throughput from indexes, to maximize storage performance. On the other hand, we realize direct NVMe-GPUs 110 stack and a CPU-GPU heterogeneous task architecture for low-latency unindexed-vector searches on streaming data. Experiments show that our approach performs with 1.5x to 3.4x bandwidth, as low as 20% latency compared to existing 110 stacks, and up to orders-of-magnitude higher vector storage throughput under concurrent RIW workloads. Further, N eos can handle real-time unindexed - vector searches with millisecond-level latency on streaming input, a capability that current vector systems lack.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Turbocharging Vector Databases using Modern SSDsJoobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do et al.VLDB 2025 · 13 citations
- Accelerating Graph Indexing for ANNS on Modern CPUsMengzhao Wang, Haotian Wu, Xiangyu Ke, Yunjun Gao et al.SIGMOD 2025 · 6 citations
- Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor SearchZhonggen Li, Xiangyu Ke, Yifan Zhu, Bocheng Yu et al.SIGMOD 2026 · 5 citations
- Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPUChangyue Liao, Mo Sun, Zihan Yang, Jun Xie et al.ICDE 2025 · 4 citations
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun et al.ICDE 2025 · 4 citations
Related papers
- VStream: A Distributed Streaming Vector Search SystemShenghao Gong, Haobo Sun, Ziquan Fang, Liu Liu et al.VLDB 2025 · 4 citations
- VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAGJunkyum Kim, Divya MahajanHPCA 2026 · 2 citations
- SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector SearchYuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun et al.VLDB 2026 · 1 citation
- NeuVSA: A Unified and Efficient Accelerator for Neural Vector SearchZiming Yuan, Lei Dai, Wen Li, Jie Zhang et al.HPCA 2025 · 1 citation
- VStore: in-storage graph based vector search acceleratorShengwen Liang, Ying Wang, Ziming Yuan, Cheng Liu et al.DAC 2022 · 20 citations
