SAM: Database Generation from Query Workloads with Supervised Autoregressive Models
Jingyi Yang, Peizhi Wu, Gao Cong, Tieying Zhang, Xiao He
摘要
With the prevalence of cloud databases, database users are increasingly reliant on the cloud database providers to manage their data. It becomes a challenge for cloud providers to benchmark different DBMS for a specific database instance without having access to the underlying data. One viable solution is to leverage a query workload, which contains a set of queries and the corresponding cardinalities, to generate a synthetic database with similar query performance. Existing methods for database generation with cardinality constraints, however, can only handle very small query workloads due to their high complexity and encounter challenges when handling join queries.
In this work, we propose SAM, a supervised deep autoregressive model-based method for database generation from query workloads. First, SAM is able to process large-scale query workloads efficiently as its complexity is linear in the size of the query workload, the number of attributes and the attribute domain size. Second, we develop algorithms to obtain unbiased samples of base relations from the deep autoregressive model and assign join keys in a way that accurately recovers the full outer join of the target database. Comprehensive experiments on real-world datasets demonstrate that SAM is able to efficiently generate a high-fidelity database that not only satisfies the input cardinality constraints, but also is close to the target database.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Controllable Tabular Data Synthesis Using Diffusion ModelsTongyu Liu, Ju Fan, Nan Tang, Guoliang Li 等SIGMOD 2024 · 被引用 13 次
- SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL WorkloadsJiale Lao, Immanuel TrummerSIGMOD 2026 · 被引用 11 次
- Privacy-Enhanced Database Synthesis for Benchmark PublishingYunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong 等VLDB 2025 · 被引用 3 次
- Automated Discovery of Test Oracles for Database Management Systems Using LLMsQiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu 等SIGMOD 2026 · 被引用 1 次
- Towards Synthesizing High-Dimensional Tabular Data with Limited SamplesZuqing Li, Junhao Gan, Jianzhong QiAAAI 2026
它引用的顶会 Paper4
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- A Unified Deep Model of Learning from both Data and Queries for Cardinality EstimationPeizhi Wu, Gao CongSIGMOD 2021 · 被引用 73 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
- Relational Data Synthesis using Generative Adversarial Networks: A Design Space ExplorationJu Fan, Tongyu Liu, Guoliang Li, Junyou Chen 等VLDB 2020
相关 Paper
- PBench: Workload Synthesizer with Real Statistics for Cloud Analytics BenchmarkingYan Zhou, Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang 等VLDB 2025 · 被引用 4 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- Graph-Conditional Flow Matching for Relational Data GenerationDavide Scassola, Sebastiano Saccani, Luca BortolussiAAAI 2026 · 被引用 3 次
- Deep Learning Models for Selectivity Estimation of Multi-Attribute QueriesShohedul Hasan, Saravanan Thirumuruganathan, Jees Augustine, Nick Koudas 等SIGMOD 2020 · 被引用 101 次
- SQLStorm: Taking Database Benchmarking into the LLM EraTobias Schmidt, Viktor Leis, Peter Boncz, Thomas NeumannVLDB 2025 · 被引用 21 次
