Approximate Partition Selection for Big-Data Workloads using Summary Statistics
Kexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula, Philip Alexander Levis
摘要
Many big-data clusters store data in large partitions that support access at a coarse, partition-level granularity. As a result, approximate query processing via row-level sampling is inefficient, often requiring reads of many partitions. In this work, we seek to answer queries quickly and approximately by reading a subset of the data partitions and combining partial answers in a weighted manner without modifying the data layout. We illustrate how to efficiently perform this query processing using a set of pre-computed summary statistics, which inform the choice of partitions and weights. We develop novel means of using the statistics to assess the similarity and importance of partitions. Our experiments on several datasets and data layouts demonstrate that to achieve the same relative error compared to uniform partition sampling, our techniques offer from 2.7x to 70x reduction in the number of partitions read, and the statistics stored per partition require fewer than 100KB.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Pre-training Summarization Models of Structured Datasets for Cardinality EstimationYao Lu, Srikanth Kandula, Arnd Christian König, Surajit ChaudhuriVLDB 2022 · 被引用 38 次
- Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query ProcessingXi Liang, Stavros Sintos, Zechao Shang, Sanjay KrishnanSIGMOD 2021 · 被引用 27 次
- S/C: Speeding up Data Materialization with Bounded MemoryZhaoheng Li, Xinyu Pi, Yongjoo ParkICDE 2023 · 被引用 7 次
- PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error GuaranteesYuxuan Zhu, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang 等SIGMOD 2025 · 被引用 3 次
- JanusAQP: Efficient Partition Tree Maintenance for Dynamic Approximate Query ProcessingXi Liang, Stavros Sintos, Sanjay KrishnanICDE 2023 · 被引用 3 次
它引用的顶会 Paper4
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke 等SIGMOD 2020 · 被引用 87 次
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 被引用 35 次
- Learning to Sample: Counting with Complex QueriesBrett Walenz, Stavros Sintos, Sudeepa Roy, Jun YangVLDB 2020 · 被引用 16 次
- CoopStore: Optimizing Precomputed Summaries for AggregationEdward Gan, Peter Bailis, Moses CharikarVLDB 2020
相关 Paper
- Salvaging failing and straggling queriesBruhathi Sundarmurthy, Harshad Deshmukh, Paris Koutris, Jeffrey F. NaughtonICDE 2022
- ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data GenerationRong Gu, Han Li, Haipeng Dai, Wenjie Huang 等VLDB 2023 · 被引用 9 次
- FAAQP: Fast and Accurate Approximate Query Processing based on Bitmap-augmented Sum-Product NetworkHanbing Zhang, Yinan Jing, Zhenying He, Kai Zhang 等SIGMOD 2025
- Random Sampling for Group-By QueriesTrong Duc Nguyen, Ming-Hung Shih, Sai Sree Parvathaneni, Bojian Xu 等ICDE 2020 · 被引用 12 次
- PPQ-Trajectory: Spatio-temporal Quantization for Querying in Large Trajectory RepositoriesShuang Wang, Hakan FerhatosmanogluVLDB 2021 · 被引用 12 次
