Toward Drift-Aware Database Benchmarking
Guanli Liu, Renata Borovica-Gajic
Abstract
Data and workload drift are critical to evaluating core database components such as caching, cardinality estimation, indexing, and query optimization, especially as AI-driven techniques increasingly permeate database systems. However, existing benchmarks remain largely static, offering little support for modeling drifts. This limitation arises from the absence of a shared vocabulary and practical tools for specifying and generating drift in both data and workloads.
Guided by this vision of making drift a first-class concept, we propose a taxonomy of data and workload drift and design DriftSpec , a declarative specification that makes these drifts executable. Building on this, we present DriftBench , which instantiates DriftSpec to generate controlled drifts and enable drift-aware benchmarking. Together, the taxonomy, DriftSpec, and DriftBench form a first step toward a standardized, executable language for studying how data and workload evolution influence database behavior. They shift benchmarking from static, one-off tests to controlled, continuous evaluation under drift.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ea83adf-9077-475f-b564-dece219e1c40Builds on22
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 180 citations
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 178 citations
- Balsa: Learning a Query Optimizer Without Expert DemonstrationsZongheng Yang, Wei-Lin Chiang, Sifei Luan, Gautam Mittal et al.SIGMOD 2022 · 99 citations
Related papers
- NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling: [Experiments & Analysis]Zhanhao Zhao, Haotian Gao, Naili Xing, Lingze Zeng et al.SIGMOD 2026
- Are Learned DBMS Components Robust to Workload Drift?: [Experiments & Analysis]Zizhong Meng, Gao Cong, Siqiang LuoSIGMOD 2026
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 57 citations
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 62 citations
- HyBench: A New Benchmark for HTAP DatabasesChao Zhang, Guoliang Li, Tao LvVLDB 2024 · 28 citations
