VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis]
Xiang Zhang, Chao Zhang, Ju Fan, Guoliang Li, Xiaoyong Du
Abstract
As vector databases become increasingly critical to supporting large language models (LLMs), efficient and effective vector data retrieval has become imperative. Filtered vector search , which retrieves relevant vectors under scalar filter constraints, has recently attracted significant attention. However, existing benchmarks for evaluating filtered vector search have three major limitations. First, they rely on simple, low-dimensional datasets that fail to reflect real-world scenarios involving thousands of dimensions and millions of records. Second, they use random filter values, resulting in trivial query plans that do not adequately challenge modern vector database optimizers. Third, they lack a unified evaluation metric to comprehensively quantify end-to-end performance. In this work, we propose VecBench, a controllable benchmark for holistic evaluation of filtered vector search. Designing such a benchmark poses three key challenges: (1) generating high-dimensional and large-scale vector data while preserving the original similarity distribution, (2) constructing representative benchmark queries that capture diverse filtered search strategies, and (3) developing a unified metric that enables fair and comprehensive comparisons under both static and dynamic scenarios. VecBench addresses these challenges through controllable data generation, adjustable query synthesis, and end-to-end holistic benchmarking. First, it can generate massive vector data with flexible dimensionality and scale while preserving the original distribution with a provable error bound. Second, it supports fine-grained workload control by varying selectivity and filter correlation, enabling systematic stress testing of different filtered search strategies. Third, it introduces a unified evaluation framework covering six phases, including concurrent query processing and dynamic update scenarios. To demonstrate the effectiveness of VecBench and to shed light on the strengths and weaknesses of different filtered vector search methods, we conduct extensive experiments on ten representative approaches across four popular vector databases.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fa781d50-734f-4731-9db6-9d9b5b86eb38Related papers
- Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL QueriesZhengren Wang, Dongwen Yao, Bozhou Li, Dongsheng Ma et al.ICDE 2026 · 1 citation
- BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector DatabasesGuoxin Kang, Zhongxin Ge, Jingpei Hu, Xueya Zhang et al.VLDB 2025 · 4 citations
- An Experimental Evaluation of Hybrid Querying on VectorsJiaxu Zhu, Jiayu Yuan, Kaiwen Yang, Xiaobao Chen et al.VLDB 2026 · 2 citations
- Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views: [Experiments & Analysis]Tingyang Chen, Cong Fu, Jiahua Wu, Haotian Wu et al.SIGMOD 2026 · 5 citations
- VecFlow: A High-Performance Vector Data Management System for Filtered-Search on GPUsJingyi Xi, Chenghao Mo, Ben Karsin, Artem M. Chirkin et al.SIGMOD 2026 · 1 citation
