Workload-Aware Incremental Reclustering in Cloud Data Warehouses
Yipeng Liu, Renfei Zhou, Jiaqi Yan, Huanchen Zhang
Abstract
Modern cloud data warehouses store data in micro-partitions and rely on metadata (e.g., zonemaps) for efficient data pruning during query processing. Maintaining data clustering in a large-scale table is crucial for effective data pruning. Existing automatic clustering approaches lack the flexibility required in dynamic cloud environments with continuous data ingestion and evolving workloads. This paper advocates a clean separation between reclustering policy and clustering-key selection. We introduce the concept of boundary micro-partitions that sit on the boundary of query ranges. We then present WAIR, a workload-aware algorithm to identify and recluster only boundary micro-partitions most critical for pruning efficiency. WAIR achieves near-optimal (with respect to fully sorted table layouts) query performance but incurs significantly lower reclustering cost with a theoretical upper bound. We further implement the algorithm into a prototype reclustering service and evaluate on standard benchmarks (TPC-H, DSB) and a real-world workload. Results show that WAIR improves query performance and reduces the overall cost compared to existing solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cd77cc2-be13-4ee8-a2c8-000810b3f3ecBuilds on6
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong et al.NSDI 2020 · 142 citations
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 62 citations
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang et al.SIGMOD 2021 · 37 citations
Related papers
- Dynamic Data Layout Optimization with Worst-Case GuaranteesKexin Rong, Paul Liu, Sarah Ashok Sonje, Moses CharikarICDE 2024 · 1 citation
- Fast and Effective Distribution-Key Recommendation for Amazon RedshiftPanos Parchas, Yonatan Naamad, Peter Van Bouwel, Christos Faloutsos et al.VLDB 2020 · 25 citations
- PAW: Data Partitioning Meets Workload VarianceZhe Li, Man Lung Yiu, Tsz Nam ChanICDE 2022 · 6 citations
- S-CDA: A Smart Cloud Disk Allocation Approach in Cloud Block Storage SystemHua Wang, Yang Yang, Ping Huang, Yu Zhang et al.DAC 2020 · 4 citations
- Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion PipelinesPatrick Hansert, Sebastian MichelVLDB 2024 · 2 citations
