High-Ratio Compression for Machine-Generated Data
Jiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng, Chuan Xiao, Wei Jia, Yue Li, Qinhui Sun, Wenjie Zhang, Xuemin Lin
摘要
Machine-generated data is rapidly growing and poses challenges for data-intensive systems, especially as the growth of data outpaces the growth of storage space. To cope with the storage issue, compression plays a critical role in storage engines, particularly for data-intensive applications, where a high compression ratio and efficient random access are essential. However, existing compression techniques tend to focus on general-purpose and data block approaches, but overlook the inherent structure of machine-generated data and hence result in low compression ratios or limited lookup efficiency. To address these limitations, we introduce the Pattern-Based Compression (PBC) algorithm, which specifically targets patterns in machine-generated data to achieve Pareto-optimality in most cases. Unlike traditional data block-based methods, PBC compresses data on a per-record basis, facilitating rapid random access. Our experimental evaluation demonstrates that PBC, on average, achieves a compression ratio twice as high as the state-of-the-art techniques while maintaining competitive compression and decompression speeds. We also integrate PBC to a production database system and achieve improvements on both comparison ratio and throughput.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 被引用 9 次
- SINDI: An Efficient Index for Sparse Vector Approximate Maximum Inner Product SearchRuoxuan Li, Xiaoyao Zhong, Jiabao Jin, Peng Cheng 等ICDE 2026 · 被引用 3 次
- Approaching Shannon Bound with Lossless LLM Weight CompressionHongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong 等ISCA 2026 · 被引用 2 次
- Morphing-based Compression for Data-centric ML PipelinesSebastian Baunsgaard, Matthias BoehmVLDB 2026
- LogLite: Lightweight Plug-and-Play Streaming Log CompressionBenzhao Tang, Shiyu Yang, Zhitao Shen, Wenjie Zhang 等VLDB 2025
它引用的顶会 Paper6
- On the Feasibility of Parser-based Log Compression in Large-Scale Cloud SystemsJunyu Wei, Guangyan Zhang, Yang Wang, Zhiwei Liu 等FAST 2021 · 被引用 34 次
- Order-Preserving Key Compression for In-Memory Search TreesHuanchen Zhang, Xiaoxuan Liu, David G. Andersen, Michael Kaminsky 等SIGMOD 2020 · 被引用 31 次
- DeepSqueeze: Deep Semantic Compression for Tabular DataAmir Ilkhechi, Andrew Crotty, Alex Galakatos, Yicong Mao 等SIGMOD 2020 · 被引用 30 次
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 28 次
- PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar StorageHao Jiang, Chunwei Liu, Qi Jin, John Paparrizos 等VLDB 2020
相关 Paper
- A Cost-Effective and Decompression-Transparent Compressor for OLTP-Oriented DatabasesHao Hu, Qiyang Zheng, Xiangyu Zou, Lisha Qin 等ICDE 2025 · 被引用 4 次
- Exploiting Inter-block Entropy to Enhance the Compressibility of Blocks with Diverse DataJinkwon Kim, Mincheol Kang, Jeongkyu Hong, Soontae KimHPCA 2022 · 被引用 5 次
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien 等SIGMOD 2021 · 被引用 45 次
- Pattern-Guided File Compression with User-Experience Enhancement for Log-Structured File System on Mobile DevicesCheng Ji, Li-Pin Chang, Riwei Pan, Chao Wu 等FAST 2021 · 被引用 31 次
- PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics DatabaseHui Sun, Yanfeng Ding, Liping Yi, Huidong Ma 等KDD 2025
