LogBase: A Large-Scale Benchmark for Semantic Log Parsing
Chenbo Zhang, Wenying Xu, Jinbu Liu, Lu Zhang, Guiyang Liu, Jihong Guan, Qi Zhou, Shuigeng Zhou
摘要
Logs generated by large-scale software systems contain a huge amount of useful information. As the first step of automated log analysis, log parsing has been extensively studied. General log parsing techniques focus on identifying static templates from raw logs, but overlook the more important semantics implied in dynamic log parameters. With the popularity of Artificial Intelligence for IT Operations (AIOps), traditional log parsing methods no longer meet the requirements of various downstream tasks. Researchers are now exploring the next generation of log parsing techniques, i.e., semantic log parsing , to identify both log templates and semantics in log parameters. However, the absence of semantic annotations in existing datasets hinders the training and evaluation of semantic log parsers, thereby stalling the progress of semantic log parsing. To fill this gap and advance the field of semantic log parsing, we construct LogBase , the first semantic log parsing benchmark dataset. LogBase consists of logs from 130 popular open-source projects, containing 85,300 semantically annotated log templates, surpassing existing datasets in both log source diversity and template richness. To build Logbase, we develop the framework GenLog for constructing semantic log parsing datasets. GenLog mines log template-parameter-context triplets from popular open-source repositories on GitHub, and uses chain-of-thought (CoT) techniques with large language models (LLMs) to generate high-quality logs. Meanwhile, GenLog employs human feedback to improve the quality of the generated data and ensure its reliability. GenLog is highly automated and cost-effective, enabling researchers to easily and efficiently construct semantic log parsing datasets. Furthermore, we also design a set of comprehensive evaluation metrics for LogBase, including general log parser metrics and the metrics specifically for semantic log parsers and LLM-based parsers. With LogBase, we extensively evaluate 15 existing log parsers, revealing their true performance in complex scenarios. We believe that this work provides researchers with valuable data, reliable tools, and insightful findings to support and guide the future research of semantic log parsing.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- SemParser: A Semantic Parser for Log AnalyticsYintong Huo, Yuxin Su, Cheryl Lee, Michael R. LyuICSE 2023 · 被引用 51 次
- A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?Zhihan Jiang, Jinyang Liu, Junjie Huang, Yichen Li 等ISSTA 2024 · 被引用 52 次
- Demonstration-Free: Towards More Practical Log Parsing with Large Language ModelsYi Xiao, Van-Hoang Le, Hongyu ZhangASE 2024 · 被引用 9 次
- DivLog: Log Parsing with Prompt Enhanced In-Context LearningJunjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang 等ICSE 2024 · 被引用 54 次
- MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context LearningJianbo Yu, Yixuan Li, Hai Xu, Kang Xu 等AAAI 2026
