CANDOR-Bench: Benchmarking In-Memory Continuous ANNS under Dynamic Open-World Streams [Experiments & Analysis]
Mingqi Wang, Junyao Dong, Zhuoyan Wu, Jun Liu, Ruicheng Zhang, Jianjun Zhao, Ruipeng Wan, Xinyan Lei, Shuhao Zhang, Bolong Zheng, Haikun Liu, Xiaofei Liao, Hai Jin
摘要
Continuous Approximate Nearest Neighbor Search (ANNS) over real-time vector data streams is an increasingly critical yet underexplored problem. In open-world settings—where data distributions shift, noise accumulates, and concurrent access is common—existing ANNS algorithms, originally designed for static or simplified streaming scenarios, struggle to balance ingestion latency, retrieval quality, and update efficiency. While benchmarks such as ANN-Benchmarks and Big-ANN-Benchmarks have standardized evaluation in static or large-scale settings, they fail to capture the nuanced, high-churn dynamics of real-world streams. We introduce CANDOR-Bench ( C ontinuous A pproximate N earest neighbor search under D ynamic O pen-wo R ld Streams, a benchmarking framework built on Big-ANN-Benchmark to evaluate in-memory ANNS under dynamic, open-world conditions. CANDOR-Bench supports high-frequency ingestion (up to hundreds of thousands of vectors per second), adaptive drift modeling (including modality shifts), stochastic noise injection, and concurrent query-update execution—all without requiring modifications to algorithm code. Across 12 datasets and 19 representative ANNS algorithms, our evaluation reveals that no single ANNS algorithm consistently delivers high recall, throughput, and update efficiency across dynamic open-world scenarios, which challenges assumptions drawn from static benchmarks. This variability reflects deeper trade-offs inherent to streaming settings. For example, smaller update batches improve data freshness but can introduce higher insertion overhead and reduce accuracy. We further observe that throughput in concurrent settings is often constrained by insertion overhead rather than query latency, which highlights a mismatch between streaming workloads and designs originally tuned for offline construction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 被引用 354 次
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 被引用 128 次
- SONG: Approximate Nearest Neighbor Search on GPUWeijie Zhao, Shulong Tan, Ping LiICDE 2020 · 被引用 103 次
- Towards Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional SpacesXi Zhao, Yao Tian, Kai Huang, Bolong Zheng 等VLDB 2023 · 被引用 88 次
相关 Paper
- CONDA: A Connectivity-Aware Dynamic Index for Approximate Nearest Neighbor Search over Evolving DataDarae Lee, Min-Soo KimVLDB 2026
- SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector SearchYuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun 等VLDB 2026 · 被引用 1 次
- OEBench: Investigating Open Environment Challenges in Real-World Relational Data StreamsYiqun Diao, Yutong Yang, Qinbin Li, Bingsheng He 等VLDB 2024 · 被引用 5 次
- High-Throughput, Cost-Effective Billion-Scale Vector Search with a Single GPUHaodi Jiang, Hao Guo, Minhui Xie, Jiwu Shu 等SIGMOD 2026
- PANNS: Enhancing Graph-based Approximate Nearest Neighbor Search through Recency-aware Construction and Parameterized SearchXizhe Yin, Chao Gao, Zhijia Zhao, Rajiv GuptaPPoPP 2025 · 被引用 5 次
