A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?
Zhihan Jiang, Jinyang Liu, Junjie Huang, Yichen Li, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Jieming Zhu, Michael R. Lyu
摘要
Log data have facilitated various tasks of software development and maintenance, such as testing, debugging and diagnosing. Due to the unstructured nature of logs, log parsing is typically required to transform log messages into structured data for automated log analysis. Given the abundance of log parsers that employ various techniques, evaluating these tools to comprehend their characteristics and performance becomes imperative. Loghub serves as a commonly used dataset for benchmarking log parsers, but it suffers from limited scale and representativeness, posing significant challenges for studies to comprehensively evaluate existing log parsers or develop new methods. This limitation is particularly pronounced when assessing these log parsers for production use. To address these limitations, we provide a new collection of annotated log datasets, denoted Loghub-2.0, which can better reflect the characteristics of log data in real-world software systems. Loghub-2.0 comprises 14 datasets with an average of 3.6 million log lines in each dataset. Based on Loghub-2.0, we conduct a thorough re-evaluation of 15 state-of-the-art log parsers in a more rigorous and practical setting. Particularly, we introduce a new evaluation metric to mitigate the sensitivity of existing metrics to imbalanced data distributions. We are also the first to investigate the granular performance of log parsers on logs that represent rare system events, offering in-depth details for software diagnosis. Accurately parsing such logs is essential, yet it remains a challenge. We believe this work could shed light on the evaluation and design of log parsers in practical settings, thereby facilitating their deployment in production systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- No More Labelled Examples? An Unsupervised Log Parser with LLMsJunjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. LyuFSE 2025 · 被引用 13 次
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language ModelsZeyang Ma, Dong Jae Kim, Tse-Hsun Peter ChenICSE 2025 · 被引用 7 次
- COCA: Generative Root Cause Analysis for Distributed Systems with Code KnowledgeYichen Li, Yulun Wu, Jinyang Liu, Zhihan Jiang 等ICSE 2025 · 被引用 6 次
- Small Is Beautiful: A Practical and Efficient Log Parsing FrameworkMinxing Wang, Yintong HuoFSE 2026 · 被引用 1 次
- CAShift: Benchmarking Log-Based Cloud Attack Detection under Normality ShiftJiongchi Yu, Xiaofei Xie, Qiang Hu, Bowen Zhang 等FSE 2025 · 被引用 1 次
它引用的顶会 Paper11
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Log-based Anomaly Detection with Deep Learning: How Far Are We?Van-Hoang Le, Hongyu ZhangICSE 2022 · 被引用 212 次
- UniParser: A Unified Log Parser for Heterogeneous Log DataYudong Liu, Xu Zhang, Shilin He, Hongyu Zhang 等WWW 2022 · 被引用 148 次
- Log Parsing with Prompt-based Few-shot LearningVan-Hoang Le, Hongyu ZhangICSE 2023 · 被引用 98 次
- Guidelines for Assessing the Accuracy of Log Message Template Identification TechniquesZanis Ali Khan, Donghwan Shin, Domenico Bianculli, Lionel C. BriandICSE 2022 · 被引用 77 次
相关 Paper
- LogParser-LLM: Advancing Efficient Log Parsing with Large Language ModelsAoxiao Zhong, Dengyao Mo, Guiyang Liu, Jinbu Liu 等KDD 2024 · 被引用 41 次
- LogBase: A Large-Scale Benchmark for Semantic Log ParsingChenbo Zhang, Wenying Xu, Jinbu Liu, Lu Zhang 等ISSTA 2025 · 被引用 5 次
- SLGParser: Practical and Efficient Label-Free Log Parsing Using Large Language ModelsYibing Hu, Cong Wang, Lixin Zhao, Aimin YuICDE 2026
- SPINE: a scalable log parser with feedback guidanceXuheng Wang, Xu Zhang, Liqun Li, Shilin He 等FSE 2022 · 被引用 48 次
- Demonstration-Free: Towards More Practical Log Parsing with Large Language ModelsYi Xiao, Van-Hoang Le, Hongyu ZhangASE 2024 · 被引用 9 次
