VDTuner: Automated Performance Tuning for Vector Data Management Systems
Tiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li, Jianguo Li, Gang Wang, Xiaoguang Liu
Abstract
Vector data management systems (VDMSs) have become an indispensable cornerstone in large-scale information retrieval and machine learning systems like large language models. To enhance the efficiency and flexibility of similarity search, VDMS exposes many tunable index parameters and system parameters for users to specify. However, due to the inherent characteristics of VDMS, automatic performance tuning for VDMS faces several critical challenges, which cannot be well addressed by the existing auto-tuning methods. In this paper, we introduce VDTuner, a learning-based automatic performance tuning framework for VDMS, leveraging multi-objective Bayesian optimization. VDTuner overcomes the challenges associated with VDMS by efficiently exploring a complex multi-dimensional parameter space without requiring any prior knowledge. Moreover, it is able to achieve a good balance between search speed and recall rate, delivering an optimal configuration. Extensive evaluations demonstrate that VDTuner can markedly improve VDMS performance (14.12% in search speed and 186.38 % in recall rate) compared with default setting, and is more efficient compared with state-of-the-art baselines (up to 3.57 x faster in terms of tuning time). In addition, VDTuner is scalable to specific user preference and cost-aware optimization objective. VDTuner is available online at https://github.com/tiannuo-yanWVDTuner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DARTH: Declarative Recall Through Early Termination for Approximate Nearest Neighbor SearchManos Chatzakis, Yannis Papakonstantinou, Themis PalpanasSIGMOD 2026 · 10 citations
- UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM ArchitectureSitian Chen, Amelie Chi Zhou, Yucheng Shi, Yusen Li et al.SC 2025 · 8 citations
- DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMsMingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang et al.SC 2025 · 5 citations
- PGTuner: An Efficient Framework for Automatic and Transferable Configuration Tuning of Proximity GraphsHao Duan, Yitong Song, Bin Yao, Anqi LiangSIGMOD 2026 · 4 citations
- Sparse Neighborhood Graph-Based Approximate Nearest Neighbor Search Revisited: Theoretical Analysis and OptimizationXinran Ma, Zhaoqi Zhou, Chuan Zhou, Zaijiu Shang et al.VLDB 2026 · 1 citation
Builds on10
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian OptimizationSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2020 · 428 citations
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin et al.SIGMOD 2021 · 113 citations
- An Inquiry into Machine Learning-based Automatic Configuration Tuning Services on Real-World Database Management SystemsDana Van Aken, Dongsheng Yang, Sebastien Brillard, Ari Fiorino et al.VLDB 2021 · 108 citations
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental EvaluationXinyi Zhang, Zhuo Chang, Yang Li, Hong Wu et al.VLDB 2022 · 88 citations
- CGPTuner: a Contextual Gaussian Process Bandit Approach for the Automatic Tuning of IT Configurations Under Varying Workload ConditionsStefano Cereda, Stefano Valladares, Paolo Cremonesi, Stefano DoniVLDB 2021 · 74 citations
Related papers
- SCOOT: SLO-Oriented Performance Tuning for LLM Inference EnginesKe Cheng, Zhi Wang, Wen Hu, Tiannuo Yang et al.WWW 2025 · 13 citations
- MINT: Multi-Vector Search Index TuningJiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein et al.ICDE 2026 · 1 citation
- VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis]Xiang Zhang, Chao Zhang, Ju Fan, Guoliang Li et al.SIGMOD 2026 · 5 citations
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 7 citations
- MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided ExplorationZihan Yan, Rui Xi, Mengshu HouSIGMOD 2026 · 3 citations
