Unlocking the Power of Numbers: Log Compression via Numeric Token Parsing
Siyu Yu, Yifan Wu, Ying Li, Pinjia He
摘要
Parser-based log compressors have been widely explored in recent years because the explosive growth of log volumes makes the compression performance of general-purpose compressors unsatisfactory. These parser-based compressors preprocess logs by grouping the logs based on the parsing result and then feed the preprocessed files into a general-purpose compressor. However, parser-based compressors have their limitations. First, the goals of parsing and compression are misaligned, so the inherent characteristics of logs were not fully utilized. In addition, the performance of parser-based compressors depends on the sample logs and thus it is very unstable. Moreover, parser-based compressors often incur a long processing time. To address these limitations, we propose Denum, a simple, general log compressor with high compression ratio and speed. The core insight is that a majority of the tokens in logs are numeric tokens (i.e. pure numbers, tokens with only numbers and special characters, and numeric variables) and effective compression of them is critical for log compression. Specifically, Denum contains a Numeric Token Parsing module, which extracts all numeric tokens and applies tailored processing methods (e.g. store the differences of incremental numbers like timestamps), and a String Processing module, which processes the remaining log content without numbers. The processed files of the two modules are then fed as input to a general-purpose compressor and it outputs the final compression results. Denum has been evaluated on 16 log datasets and it achieves an 8.7% -- 434.7% higher average compression ratio and 2.6× -- 37.7× faster average compression speed (i.e. 26.2 MB/S) compared to the baselines. Moreover, integrating Denum's Numeric Token Parsing module into existing log compressors can provide a 11.8% improvement in their average compression ratio and achieve 37% faster average compression speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LogNexus: Effective Log Compression via Unified Redundancy EncodingYang Liu, Kaiming Zhang, Zhuangbin Chen, Zibin ZhengISSTA 2026
- LogFold: Compressing Logs with Structured Tokens and Hybrid EncodingShiwen Shan, Yintong Huo, Hongzhan Zhong, Zhining Wang 等ICSE 2026
- DNSLogzip: A Novel Approach to Fast and High-Ratio Compression for DNS LogsYunwei Dai, Guyue Liu, Tao Huang, Shuo Wang 等SIGCOMM 2025
它引用的顶会 Paper10
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- UniParser: A Unified Log Parser for Heterogeneous Log DataYudong Liu, Xu Zhang, Shilin He, Hongyu Zhang 等WWW 2022 · 被引用 148 次
- Log Parsing with Prompt-based Few-shot LearningVan-Hoang Le, Hongyu ZhangICSE 2023 · 被引用 98 次
- SemParser: A Semantic Parser for Log AnalyticsYintong Huo, Yuxin Su, Cheryl Lee, Michael R. LyuICSE 2023 · 被引用 51 次
- SPINE: a scalable log parser with feedback guidanceXuheng Wang, Xu Zhang, Liqun Li, Shilin He 等FSE 2022 · 被引用 48 次
相关 Paper
- On the Feasibility of Parser-based Log Compression in Large-Scale Cloud SystemsJunyu Wei, Guangyan Zhang, Yang Wang, Zhiwei Liu 等FAST 2021 · 被引用 34 次
- LogDelta: Differential Encoding for Log DataSongze Li, Shaoxu Song, Zhitao ShenICDE 2026
- SLGParser: Practical and Efficient Label-Free Log Parsing Using Large Language ModelsYibing Hu, Cong Wang, Lixin Zhao, Aimin YuICDE 2026
- LogShrink: Effective Log Compression by Leveraging Commonality and Variability of Log DataXiaoyun Li, Hongyu Zhang, Van-Hoang Le, Pengfei ChenICSE 2024 · 被引用 22 次
- Small Is Beautiful: A Practical and Efficient Log Parsing FrameworkMinxing Wang, Yintong HuoFSE 2026 · 被引用 1 次
