Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion Pipelines
Patrick Hansert, Sebastian Michel
Abstract
Data Lakes deployed in the cloud are a go-to solution for enterprise data storage. While the pay-as-you-go cost model allows flexible resource allocation and billing, it mandates an efficient use of resources like CPU hours, network traffic, and used storage. The distributed nature of cloud environments necessitates partitioning the data and processing these partitions separately. In this work, we put forward a practical solution to improve the efficiency of compression algorithms on Dremel-encoded data by clustering similarly structured nested data at ingestion time, such that compressible partitions can be created. We propose a clustering approach inspired by decision trees that outpaces even the naive partition-then-sort approach by up to factor 17.44 while also boosting the compression by up to factor 2. We further show that when sorting the individual buckets, a compression boost that is competitive with the well-established increasing-cardinality heuristic can be achieved, but at a lower ingestion time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caff937a-df04-4b3d-8685-4dc7f6890c8eBuilds on9
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang et al.SIGMOD 2021 · 37 citations
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 28 citations
- Proteus: Autonomous Adaptive Storage for Mixed WorkloadsMichael Abebe, Horatiu Lazu, Khuzaima DaudjeeSIGMOD 2022 · 20 citations
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt et al.SIGMOD 2021 · 18 citations
- LogGrep: Fast and Cheap Cloud Log Storage by Exploiting both Static and Runtime PatternsJunyu Wei, Guangyan Zhang, Junchao Chen, Yang Wang et al.EuroSys 2023 · 18 citations
Related papers
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh et al.ICDE 2023 · 8 citations
- Workload-Aware Incremental Reclustering in Cloud Data WarehousesYipeng Liu, Renfei Zhou, Jiaqi Yan, Huanchen ZhangSIGMOD 2026 · 1 citation
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
- Robust and Budget-Constrained Encoding Configurations for In-Memory Database SystemsMartin BoissierVLDB 2022 · 17 citations
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng et al.SIGMOD 2024 · 7 citations
