DeepDB: Learn from Data, not from Queries!
Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina, Kristian Kersting, Carsten Binnig
Abstract
The typical approach for learned DBMS components is to capture the behavior by running a representative set of queries and use the observations to train a machine learning model. This workload-driven approach, however, has two major downsides. First, collecting the training data can be very expensive, since all queries need to be executed on potentially large databases. Second, training data has to be recollected when the workload and the data changes. To overcome these limitations, we take a different route: we propose to learn a pure data-driven model that can be used for different tasks such as query answering or cardinality estimation. This data-driven model also supports ad-hoc queries and updates of the data without the need of full retraining when the workload or data changes. Indeed, one may now expect that this comes at a price of lower accuracy since workload-driven models can make use of more information. However, this is not the case. The results of our empirical evaluation demonstrate that our data-driven approach not only provides better accuracy than state-ofthe-art learned components but also generalizes better to unseen queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 613f0c6a-b482-4d35-ab55-cf5cfd0e213bCited by top-tier papers56
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu et al.VLDB 2022 · 169 citations
- FLAT: Fast, Lightweight and Accurate Method for Cardinality EstimationRong Zhu, Ziniu Wu, Yuxing Han, Kai Zeng et al.VLDB 2021 · 120 citations
- Lero: A Learning-to-Rank Query OptimizerRong Zhu, Wei Chen, Bolin Ding, Xingguang Chen et al.VLDB 2023 · 102 citations
- FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory SystemsPengfei Li, Yu Hua, Jingnan Jia, Pengfei ZuoVLDB 2022 · 97 citations
- Zero-Shot Cost Models for Out-of-the-box Learned Cost PredictionBenjamin Hilprecht, Carsten BinnigVLDB 2022 · 90 citations
Builds on2
Related papers
- Are Learned DBMS Components Robust to Workload Drift?: [Experiments & Analysis]Zizhong Meng, Gao Cong, Siqiang LuoSIGMOD 2026
- Robust Query Driven Cardinality Estimation under Changing WorkloadsParimarjan Negi, Ziniu Wu, Andreas Kipf, Nesime Tatbul et al.VLDB 2023 · 88 citations
- Data-Agnostic Cardinality Learning from Imperfect WorkloadsPeizhi Wu, Rong Kang, Tieying Zhang, Jianjun Chen et al.VLDB 2025 · 1 citation
- Expand your Training Limits! Generating Training Data for ML-based Data ManagementFrancesco Ventura, Zoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Volker MarklSIGMOD 2021 · 19 citations
- Are We Ready For Learned Cardinality Estimation?Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang et al.VLDB 2021 · 156 citations
