Conditional Generative Model Based Predicate-Aware Query Approximation
Nikhil Sheoran, Subrata Mitra, Vibhor Porwal, Siddharth Ghetia, Jatin Varshney, Tung Mai, Anup B. Rao, Vikas Maddukuri
摘要
The goal of Approximate Query Processing (AQP) is to provide very fast but "accurate enough" results for costly aggregate queries thereby improving user experience in interactive exploration of large datasets. Recently proposed Machine-Learning-based AQP techniques can provide very low latency as query execution only involves model inference as compared to traditional query processing on database clusters. However, with increase in the number of filtering predicates (WHERE clauses), the approximation error significantly increases for these methods. Analysts often use queries with a large number of predicates for insights discovery. Thus, maintaining low approximation error is important to prevent analysts from drawing misleading conclusions. In this paper, we propose ELECTRA 1 , a predicate-aware AQP system that can answer analytics-style queries with a large number of predicates with much smaller approximation errors. ELEC-TRA uses a conditional generative model that learns the conditional distribution of the data and at run-time generates a small (≈ 1000 rows) but representative sample, on which the query is executed to compute the approximate result. Our evaluations with four different baselines on three real-world datasets show that ELECTRA provides lower AQP error for large number of predicates compared to baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Approximate Caching for Efficiently Serving Text-to-Image Diffusion ModelsShubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam 等NSDI 2024 · 被引用 44 次
- SEIDEN: Revisiting Query Processing in Video Database SystemsJaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra 等VLDB 2023 · 被引用 24 次
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang 等VLDB 2023 · 被引用 9 次
- A Step Toward Deep Online AggregationNikhil Sheoran, Supawit Chockchowwat, Arav Chheda, Suwen Wang 等SIGMOD 2023 · 被引用 7 次
- PairwiseHist: Fast, Accurate, and Space-Efficient Approximate Query Processing with Data CompressionAaron Hurst, Daniel E. Lucani, Qi ZhangVLDB 2024 · 被引用 5 次
它引用的顶会 Paper6
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
- Deep Learning Models for Selectivity Estimation of Multi-Attribute QueriesShohedul Hasan, Saravanan Thirumuruganathan, Jees Augustine, Nick Koudas 等SIGMOD 2020 · 被引用 101 次
- Approximate Query Processing for Data Exploration using Deep Generative ModelsSaravanan Thirumuruganathan, Shohedul Hasan, Nick Koudas, Gautam DasICDE 2020 · 被引用 54 次
相关 Paper
- FAAQP: Fast and Accurate Approximate Query Processing based on Bitmap-augmented Sum-Product NetworkHanbing Zhang, Yinan Jing, Zhenying He, Kai Zhang 等SIGMOD 2025
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto 等VLDB 2021 · 被引用 34 次
- LAQy: Efficient and Reusable Query Approximations via Lazy SamplingViktor Sanca, Periklis Chrysogelos, Anastasia AilamakiSIGMOD 2023 · 被引用 5 次
- NeuroSketch: Fast and Approximate Evaluation of Range Aggregate Queries with Neural NetworksSepanta Zeighami, Cyrus Shahabi, Vatsal SharanSIGMOD 2023 · 被引用 9 次
- ThalamusDB: Approximate Query Processing on Multi-Modal DataSaehan Jo, Immanuel TrummerSIGMOD 2024 · 被引用 11 次
