RED-ANNS: A RDMA-Enabled Distributed Framework for Graph-Based Approximate Nearest Neighbor Search
Yue Chen, Kai Zhang, Sipeng Chen, Shihai Xiao, Xiaomin Zou, Ren Ren, Yinan Jing, X. Sean Wang, Li Cao, Mingxiang Wan
Abstract
Unstructured data, such as text and images, are converted into high-dimensional vectors to capture their semantics for effective data retrieval. Approximate Nearest Neighbor Search (ANNS) over these vectors has become a fundamental technique in many domains, including retrieval-augmented generation and recommendation systems. With an ever-increasing volume of data, existing distributed solutions typically segment data across multiple machine nodes, handling query processing in a MapReduce-style approach. However, this approach suffers from reduced indexing efficiency and increased computational overhead, resulting in limited performance enhancement despite investing several times more resources. In this work, we propose RED-ANNS, a distributed ANNS approach on an RDMA network. The core idea is to maintain a logically full graph across a shared memory space of multiple nodes and utilize Remote Direct Memory Access (RDMA) to search the distributed graph, thereby avoiding the reduction in indexing efficiency caused by segmentation. The key to making this approach effective is to address the overhead associated with remote accesses. We reduce remote access frequency through locality-aware data placement and affinity-based query scheduling, while we hide remote access latency with a dependency-relaxed best-first search algorithm. Extensive experiments demonstrate that RED-ANNS achieves a performance improvement of up to 2.5× over MapReduce-style approaches and up to 5.3× over open source vector databases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fe01f5f-fbdd-4afe-8933-a2910c278d77Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 354 citations
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li et al.NeurIPS 2021 · 219 citations
- HM-ANN: Efficient Billion-Point Nearest Neighbor Search on Heterogeneous MemoryJie Ren, Minjia Zhang, Dong LiNeurIPS 2020 · 136 citations
Related papers
- DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMsMingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang et al.SC 2025 · 5 citations
- HARMONY: A Scalable Distributed Vector Database for High-Throughput Approximate Nearest Neighbor SearchQian Xu, Feng Zhang, Chengxi Li, Lei Cao et al.SIGMOD 2026 · 7 citations
- NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingYitu Wang, Shiyu Li, Qilin Zheng, Linghao Song et al.ISCA 2024 · 26 citations
- Quake: Adaptive Indexing for Vector SearchJason Mohoney, Devesh Sarda, Mengze Tang, Shihabur Rahman Chowdhury et al.OSDI 2025 · 12 citations
- FlashANNS: GPU-Driven Asynchronous I/O Pipelining for Eliminating Storage-Compute Bottlenecks in Billion-Scale Similarity SearchYang Xiao, Mo Sun, Ziyu Song, Bing Tian et al.SIGMOD 2026 · 3 citations
