Robust Query Driven Cardinality Estimation under Changing Workloads
Parimarjan Negi, Ziniu Wu, Andreas Kipf, Nesime Tatbul, Ryan Marcus, Sam Madden, Tim Kraska, Mohammad Alizadeh
Abstract
Query driven cardinality estimation models learn from a historical log of queries. They are lightweight, having low storage requirements, fast inference and training, and are easily adaptable for any kind of query. Unfortunately, such models can suffer unpredictably bad performance under workload drift, i.e., if the query pattern or data changes. This makes them unreliable and hard to deploy. We analyze the reasons why models become unpredictable due to workload drift, and introduce modifications to the query representation and neural network training techniques to make query-driven models robust to the effects of workload drift. First, we emulate workload drift in queries involving some unseen tables or columns by randomly masking out some table or column features during training. This forces the model to make predictions with missing query information, relying more on robust features based on up-to-date DBMS statistics that are useful even when query or data drift happens. Second, we introduce join bitmaps, which extends sampling-based features to be consistent across joins using ideas from sideways information passing. Finally, we show how both of these ideas can be adapted to handle data updates.
We show significantly greater generalization than past works across different workloads and databases. For instance, a model trained with our techniques on a simple workload (JOBLight-train), with 40 k synthetically generated queries of at most 3 tables each, is able to generalize to the much more complex Join Order Benchmark, which include queries with up to 16 tables, and improve query runtimes by 2× over PostgreSQL. We show similar robustness results with data updates, and across other workloads. We discuss the situations where we expect, and see, improvements, as well as more challenging workload drift scenarios where these techniques do not improve much over PostgreSQL. However, even in the most challenging scenarios, our models never perform worse than PostgreSQL, while standard query driven models can get much worse than PostgreSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44bd4af6-2f06-4955-b50d-dfe5fa23a3d3Cited by top-tier papers32
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic WorkloadsPengfei Li, Wenqing Wei, Rong Zhu, Bolin Ding et al.VLDB 2024 · 50 citations
- Breaking It Down: An In-depth Study of Index AdvisorsWei Zhou, Chen Lin, Xuanhe Zhou, Guoliang LiVLDB 2024 · 21 citations
- Sample-Efficient Cardinality Estimation Using Geometric Deep LearningSilvan Reiner, Michael GrossniklausVLDB 2024 · 20 citations
- PRICE: A Pretrained Model for Cross-Database Cardinality EstimationTianjing Zeng, Junwei Lan, Jiahong Ma, Wenqing Wei et al.VLDB 2025 · 14 citations
- The Holon Approach for Simultaneously Tuning Multiple Components in a Self-Driving Database Management System with Machine Learning via Synthesized Proto-ActionsWilliam Zhang, Wan Shen Lim, Matthew Butrovich, Andrew PavloVLDB 2024 · 13 citations
Builds on17
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu et al.VLDB 2022 · 169 citations
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 156 citations
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
Related papers
- Are Learned DBMS Components Robust to Workload Drift?: [Experiments & Analysis]Zizhong Meng, Gao Cong, Siqiang LuoSIGMOD 2026
- Warper: Efficiently Adapting Learned Cardinality Estimators to Data and Workload DriftsBeibin Li, Yao Lu, Srikanth KandulaSIGMOD 2022 · 29 citations
- Data-Agnostic Cardinality Learning from Imperfect WorkloadsPeizhi Wu, Rong Kang, Tieying Zhang, Jianjun Chen et al.VLDB 2025 · 1 citation
- Duet: Efficient and Scalable Hybrid Neural Relation UnderstandingKaixin Zhang, Hongzhi Wang, Yabin Lu, Ziqi Li et al.ICDE 2024 · 3 citations
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
