BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases
Guoxin Kang, Zhongxin Ge, Jingpei Hu, Xueya Zhang, Lei Wang, Jianfeng Zhan
Abstract
Vector databases are designed to effectively store, organize, and retrieve high-dimensional vectors, enabling faster and more accurate querying and analysis. This study highlights that the performance of cutting-edge vector databases hinges on their proficiency in managing heterogeneous data embedding and handling compound queries. The former task revolves around converting varied data types into a cohesive vector format, while the latter involves processing multimodal or single-modal queries with precise constraints. The paper advocates for evaluating these dual tasks within an integrated benchmark framework. However, state-of-the-art vector database benchmarks overlook heterogeneous data embedding and compound queries, creating a gap in evaluating vector database performance. To address this gap, we introduce BigVectorBench, a benchmark suite designed to evaluate vector database performance. BigVectorBench contributes by defining and evaluating the embedding performance of heterogeneous data. Additionally, it abstracts compound queries, which are increasingly used in real-world applications, replacing unimodal vector searches. Our rigorous evaluations validate the two design decisions of BigVectorBench and identify performance bottlenecks of mainstream vector databases. Its source code and user manual are available from https://github.com/BenchCouncil/BigVectorBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3563cd0c-4928-4ce2-afad-f0026536ee5dCited by top-tier papers1
Ask how each one uses itBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis]Xiang Zhang, Chao Zhang, Ju Fan, Guoliang Li et al.SIGMOD 2026 · 5 citations
- MINT: Multi-Vector Search Index TuningJiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein et al.ICDE 2026 · 1 citation
- Are There Fundamental Limitations in Supporting Vector Data Management in Relational Databases? A Case Study of PostgreSQLYunan Zhang, Shige Liu, Jianguo WangICDE 2024 · 26 citations
- Text2VectorSQL: Towards a Unified Interface for Vector Search and SQL QueriesZhengren Wang, Dongwen Yao, Bozhou Li, Dongsheng Ma et al.ICDE 2026 · 1 citation
- An Experimental Evaluation of Hybrid Querying on VectorsJiaxu Zhu, Jiayu Yuan, Kaiwen Yang, Xiaobao Chen et al.VLDB 2026 · 2 citations
