High-Ratio Compression for Machine-Generated Data
Jiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng, Chuan Xiao, Wei Jia, Yue Li, Qinhui Sun, Wenjie Zhang, Xuemin Lin
Abstract
Machine-generated data is rapidly growing and poses challenges for data-intensive systems, especially as the growth of data outpaces the growth of storage space. To cope with the storage issue, compression plays a critical role in storage engines, particularly for data-intensive applications, where a high compression ratio and efficient random access are essential. However, existing compression techniques tend to focus on general-purpose and data block approaches, but overlook the inherent structure of machine-generated data and hence result in low compression ratios or limited lookup efficiency. To address these limitations, we introduce the Pattern-Based Compression (PBC) algorithm, which specifically targets patterns in machine-generated data to achieve Pareto-optimality in most cases. Unlike traditional data block-based methods, PBC compresses data on a per-record basis, facilitating rapid random access. Our experimental evaluation demonstrates that PBC, on average, achieves a compression ratio twice as high as the state-of-the-art techniques while maintaining competitive compression and decompression speeds. We also integrate PBC to a production database system and achieve improvements on both comparison ratio and throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2adcc89e-83ad-4ae7-9021-d62164ddafe6Cited by top-tier papers6
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 9 citations
- SINDI: An Efficient Index for Sparse Vector Approximate Maximum Inner Product SearchRuoxuan Li, Xiaoyao Zhong, Jiabao Jin, Peng Cheng et al.ICDE 2026 · 3 citations
- Approaching Shannon Bound with Lossless LLM Weight CompressionHongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong et al.ISCA 2026 · 2 citations
- Morphing-based Compression for Data-centric ML PipelinesSebastian Baunsgaard, Matthias BoehmVLDB 2026
- LogLite: Lightweight Plug-and-Play Streaming Log CompressionBenzhao Tang, Shiyu Yang, Zhitao Shen, Wenjie Zhang et al.VLDB 2025
Builds on6
- On the Feasibility of Parser-based Log Compression in Large-Scale Cloud SystemsJunyu Wei, Guangyan Zhang, Yang Wang, Zhiwei Liu et al.FAST 2021 · 34 citations
- Order-Preserving Key Compression for In-Memory Search TreesHuanchen Zhang, Xiaoxuan Liu, David G. Andersen, Michael Kaminsky et al.SIGMOD 2020 · 31 citations
- DeepSqueeze: Deep Semantic Compression for Tabular DataAmir Ilkhechi, Andrew Crotty, Alex Galakatos, Yicong Mao et al.SIGMOD 2020 · 30 citations
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 28 citations
- PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar StorageHao Jiang, Chunwei Liu, Qi Jin, John Paparrizos et al.VLDB 2020
Related papers
- A Cost-Effective and Decompression-Transparent Compressor for OLTP-Oriented DatabasesHao Hu, Qiyang Zheng, Xiangyu Zou, Lisha Qin et al.ICDE 2025 · 4 citations
- Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse DataJinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae KimHPCA 2022 · 5 citations
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien et al.SIGMOD 2021 · 45 citations
- Pattern-Guided File Compression with User-Experience Enhancement for Log-Structured File System on Mobile DevicesCheng Ji, Li-Pin Chang, Riwei Pan, Chao Wu et al.FAST 2021 · 31 citations
- PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics DatabaseHui Sun, Yanfeng Ding, Liping Yi, Huidong Ma et al.KDD 2025
