CANDOR-Bench: Benchmarking In-Memory Continuous ANNS under Dynamic Open-World Streams [Experiments & Analysis]
Mingqi Wang, Junyao Dong, Zhuoyan Wu, Jun Liu, Ruicheng Zhang, Jianjun Zhao, Ruipeng Wan, Xinyan Lei, Shuhao Zhang, Bolong Zheng, Haikun Liu, Xiaofei Liao, Hai Jin
Abstract
Continuous Approximate Nearest Neighbor Search (ANNS) over real-time vector data streams is an increasingly critical yet underexplored problem. In open-world settings—where data distributions shift, noise accumulates, and concurrent access is common—existing ANNS algorithms, originally designed for static or simplified streaming scenarios, struggle to balance ingestion latency, retrieval quality, and update efficiency. While benchmarks such as ANN-Benchmarks and Big-ANN-Benchmarks have standardized evaluation in static or large-scale settings, they fail to capture the nuanced, high-churn dynamics of real-world streams. We introduce CANDOR-Bench ( C ontinuous A pproximate N earest neighbor search under D ynamic O pen-wo R ld Streams, a benchmarking framework built on Big-ANN-Benchmark to evaluate in-memory ANNS under dynamic, open-world conditions. CANDOR-Bench supports high-frequency ingestion (up to hundreds of thousands of vectors per second), adaptive drift modeling (including modality shifts), stochastic noise injection, and concurrent query-update execution—all without requiring modifications to algorithm code. Across 12 datasets and 19 representative ANNS algorithms, our evaluation reveals that no single ANNS algorithm consistently delivers high recall, throughput, and update efficiency across dynamic open-world scenarios, which challenges assumptions drawn from static benchmarks. This variability reflects deeper trade-offs inherent to streaming settings. For example, smaller update batches improve data freshness but can introduce higher insertion overhead and reduce accuracy. We further observe that throughput in concurrent settings is often constrained by insertion overhead rather than query latency, which highlights a mismatch between streaming workloads and designs originally tuned for offline construction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa146c86-edcd-4482-891b-ef4d438ab186Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 354 citations
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- SONG: Approximate Nearest Neighbor Search on GPUWeijie Zhao, Shulong Tan, Ping LiICDE 2020 · 103 citations
- Towards Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional SpacesXi Zhao, Yao Tian, Kai Huang, Bolong Zheng et al.VLDB 2023 · 88 citations
Related papers
- CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving DataDarae Lee, Min-Soo KimVLDB 2026
- SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector SearchYuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun et al.VLDB 2026 · 1 citation
- OEBench: Investigating Open Environment Challenges in Real-World Relational Data StreamsYiqun Diao, Yutong Yang, Qinbin Li, Bingsheng He et al.VLDB 2024 · 5 citations
- High-Throughput, Cost-Effective Billion-Scale Vector Search with a Single GPUHaodi Jiang, Hao Guo, Minhui Xie, Jiwu Shu et al.SIGMOD 2026
- PANNS: Enhancing Graph-based Approximate Nearest Neighbor Search through Recency-aware Construction and Parameterized SearchXizhe Yin, Chao Gao, Zhijia Zhao, Rajiv GuptaPPoPP 2025 · 5 citations
