SAM: Database Generation from Query Workloads with Supervised Autoregressive Models
Jingyi Yang, Peizhi Wu, Gao Cong, Tieying Zhang, Xiao He
Abstract
With the prevalence of cloud databases, database users are increasingly reliant on the cloud database providers to manage their data. It becomes a challenge for cloud providers to benchmark different DBMS for a specific database instance without having access to the underlying data. One viable solution is to leverage a query workload, which contains a set of queries and the corresponding cardinalities, to generate a synthetic database with similar query performance. Existing methods for database generation with cardinality constraints, however, can only handle very small query workloads due to their high complexity and encounter challenges when handling join queries.
In this work, we propose SAM, a supervised deep autoregressive model-based method for database generation from query workloads. First, SAM is able to process large-scale query workloads efficiently as its complexity is linear in the size of the query workload, the number of attributes and the attribute domain size. Second, we develop algorithms to obtain unbiased samples of base relations from the deep autoregressive model and assign join keys in a way that accurately recovers the full outer join of the target database. Comprehensive experiments on real-world datasets demonstrate that SAM is able to efficiently generate a high-fidelity database that not only satisfies the input cardinality constraints, but also is close to the target database.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bafe5842-4f3b-4740-8496-d7e2b9e2dadfCited by top-tier papers5
- Controllable Tabular Data Synthesis Using Diffusion ModelsTongyu Liu, Ju Fan, Nan Tang, Guoliang Li et al.SIGMOD 2024 · 13 citations
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 11 citations
- Privacy-Enhanced Database Synthesis for Benchmark PublishingYunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong et al.VLDB 2025 · 3 citations
- Automated Discovery of Test Oracles for Database Management Systems Using LLMsQiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu et al.SIGMOD 2026 · 1 citation
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
Builds on4
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- A Unified Deep Model of Learning from both Data and Queries for Cardinality EstimationPeizhi Wu, Gao CongSIGMOD 2021 · 73 citations
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 54 citations
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen et al.VLDB 2020
Related papers
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang et al.VLDB 2025 · 4 citations
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
- Graph-Conditional Flow Matching for Relational Data GenerationDavide Scassola, Sebastiano Saccani, Luca BortolussiAAAI 2026 · 3 citations
- Deep Learning Models for Selectivity Estimation of Multi-Attribute QueriesShohedul Hasan, Saravanan Thirumuruganathan, Jees Augustine, Nick Koudas et al.SIGMOD 2020 · 101 citations
- SQLStorm: Taking Database Benchmarking into the LLM EraTobias Schmidt, Viktor Leis, Peter Boncz, Thomas NeumannVLDB 2025 · 21 citations
