Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion Pipelines
Patrick Hansert, Sebastian Michel
摘要
Data Lakes deployed in the cloud are a go-to solution for enterprise data storage. While the pay-as-you-go cost model allows flexible resource allocation and billing, it mandates an efficient use of resources like CPU hours, network traffic, and used storage. The distributed nature of cloud environments necessitates partitioning the data and processing these partitions separately. In this work, we put forward a practical solution to improve the efficiency of compression algorithms on Dremel-encoded data by clustering similarly structured nested data at ingestion time, such that compressible partitions can be created. We propose a clustering approach inspired by decision trees that outpaces even the naive partition-then-sort approach by up to factor 17.44 while also boosting the compression by up to factor 2. We further show that when sorting the individual buckets, a compression boost that is competitive with the well-established increasing-cardinality heuristic can be achieved, but at a lower ingestion time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang 等SIGMOD 2021 · 被引用 37 次
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 28 次
- Proteus: Autonomous Adaptive Storage for Mixed WorkloadsMichael Abebe, Horatiu Lazu, Khuzaima DaudjeeSIGMOD 2022 · 被引用 20 次
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt 等SIGMOD 2021 · 被引用 18 次
- LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime PatternsJunyu Wei, Guangyan Zhang, Junchao Chen, Yang Wang 等EuroSys 2023 · 被引用 18 次
相关 Paper
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh 等ICDE 2023 · 被引用 8 次
- Workload-Aware Incremental Reclustering in Cloud Data WarehousesYipeng Liu, Renfei Zhou, Jiaqi Yan, Huanchen ZhangSIGMOD 2026 · 被引用 1 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Robust and Budget-Constrained Encoding Configurations for In-Memory Database SystemsMartin BoissierVLDB 2022 · 被引用 17 次
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng 等SIGMOD 2024 · 被引用 7 次
