Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing
Xi Liang, Stavros Sintos, Zechao Shang, Sanjay Krishnan
摘要
Sample-based approximate query processing (AQP) suffers from many pitfalls such as the inability to answer very selective queries and unreliable confidence intervals when sample sizes are small. Recent research presented an intriguing solution of combining materialized, pre-computed aggregates with sampling for accurate and more reliable AQP. We explore this solution in detail in this work and propose an AQP physical design called PASS, or Precomputation-Assisted Stratified Sampling. PASS builds a tree of partial aggregates that cover different partitions of the dataset. The leaf nodes of this tree form the strata for stratified samples. Aggregate queries whose predicates align with the partitions (or unions of partitions) are exactly answered with a depth-first search, and any partial overlaps are approximated with the stratified samples. We propose an algorithm for optimally partitioning the data into such a data structure with various practical approximation techniques.
- A version of this paper has been accepted to SIGMOD'21. This document is its associated technical report. This work is mainly done when Zechao was at the University of Chicago.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ALECE: An Attention-based Learned Cardinality Estimator for SPJ Queries on Dynamic WorkloadsPengfei Li, Wenqing Wei, Rong Zhu, Bolin Ding 等VLDB 2024 · 被引用 50 次
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
- LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime PatternsJunyu Wei, Guangyan Zhang, Junchao Chen, Yang Wang 等EuroSys 2023 · 被引用 18 次
- ThalamusDB: Approximate Query Processing on Multi-Modal DataSaehan Jo, Immanuel TrummerSIGMOD 2024 · 被引用 11 次
- PairwiseHist: Fast, Accurate, and Space-Efficient Approximate Query Processing with Data CompressionAaron Hurst, Daniel E. Lucani, Qi ZhangVLDB 2024 · 被引用 5 次
它引用的顶会 Paper5
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina 等VLDB 2020 · 被引用 154 次
- Learning to Sample: Counting with Complex QueriesBrett Walenz, Stavros Sintos, Sudeepa Roy, Jun YangVLDB 2020 · 被引用 16 次
- Approximate Partition Selection for Big-Data Workloads using Summary StatisticsKexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula 等VLDB 2020 · 被引用 8 次
- CoopStore: Optimizing Precomputed Summaries for AggregationEdward Gan, Peter Bailis, Moses CharikarVLDB 2020
相关 Paper
- AB-tree: Index for Concurrent Random Sampling and UpdatesZhuoyue Zhao, Dong Xie, Feifei LiVLDB 2022 · 被引用 7 次
- Accelerating Approximate Aggregation Queries with Expensive PredicatesDaniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto 等VLDB 2021 · 被引用 34 次
- Join on Samples: A Theoretical Guide for PractitionersDawei Huang, Dong Young Yoon, Seth Pettie, Barzan MozafariVLDB 2020 · 被引用 13 次
- Rapid Approximate Aggregation with Distribution-Sensitive Interval GuaranteesStephen Macke, Maryam Aliakbarpour, Ilias Diakonikolas, Aditya G. Parameswaran 等ICDE 2021 · 被引用 3 次
- PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error GuaranteesYuxuan Zhu, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang 等SIGMOD 2025 · 被引用 3 次
