PAW: Data Partitioning Meets Workload Variance
Zhe Li, Man Lung Yiu, Tsz Nam Chan
Abstract
In distributed storage systems (e.g., HDFS, Amazon S3, Databricks), partitioning is applied on a dataset in order to enhance performance and availability. Recently, partitioning methods have been designed to optimize the query performance of partitions with respect to the historical query workload. Never-theless, in practice, future query workloads may deviate from the historical query workload, thus deteriorating the performance of existing partitioning methods. To fill this research gap, we model the variance of future query workloads from the historical query workload, then exploit this characteristic to produce partitions that perform well for future query workloads. In addition, we explore the space of irregular shaped partition regions to further optimize the query performance. Experimental results on TPC-H and real datasets show that our proposal is up to 70x more efficient than the state-of-the-art method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0bcd073-5ef9-469c-9945-b133ce2b17e9Cited by top-tier papers2
- Adaptive Indexing of Objects with Spatial ExtentFatemeh Zardbani, Nikos Mamoulis, Stratos Idreos, Panagiotis KarrasVLDB 2023 · 15 citations
- Dynamic Data Layout Optimization with Worst-Case GuaranteesKexin Rong, Paul Liu, Sarah Ashok Sonje, Moses CharikarICDE 2024 · 1 citation
Builds on2
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang et al.SIGMOD 2021 · 37 citations
Related papers
- DEPA-Delta Shifting and Distribution Shaping for Efficient Adaptive IndexingAhmad Khazaie, Holger PirkICDE 2025
- Workload-Aware Incremental Reclustering in Cloud Data WarehousesYipeng Liu, Renfei Zhou, Jiaqi Yan, Huanchen ZhangSIGMOD 2026 · 1 citation
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh et al.ICDE 2023 · 8 citations
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 3 citations
- Lachesis: Automated Partitioning for UDF-Centric AnalyticsJia Zou, Amitabh Das, Pratik Barhate, Arun Iyengar et al.VLDB 2021 · 1 citation
