LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language Models
Zeyang Ma, Dong Jae Kim, Tse-Hsun Peter Chen
摘要
Log parsing is a critical step that transforms unstructured log data into structured formats, facilitating subsequent log-based analysis. Traditional syntax-based log parsers are efficient and effective, but they often experience decreased accuracy when processing logs that deviate from the predefined rules. Recently, large language models (LLM) based log parsers have shown superior parsing accuracy. However, existing LLM-based parsers face three main challenges: 1) time-consuming and labor-intensive manual labeling for fine-tuning or in-context learning, 2) increased parsing costs due to the vast volume of log data and limited context size of LLMs, and 3) privacy risks from using commercial models like ChatGPT with sensitive log information. To overcome these limitations, this paper introduces LibreLog, an unsupervised log parsing approach that leverages open-source LLMs (i.e., Llama3-8B) to enhance privacy and reduce operational costs while achieving state-of-the-art parsing accuracy. LibreLog first groups logs with similar static text but varying dynamic variables using a fixed-depth grouping tree. It then parses logs within these groups using three components: i) similarity scoring-based retrieval augmented generation: selects diverse logs within each group based on Jaccard similarity, helping the LLM distinguish between static text and dynamic variables; ii) self-reflection: iteratively query LLMs to refine log templates to improve parsing accuracy; and iii) log template memory: stores parsed templates to reduce LLM queries for improved parsing efficiency. Our evaluation on LogHub-2.0 shows that LibreLog achieves 25% higher parsing accuracy and processes logs 2.7 times faster compared to state-of-the-art LLM-based parsers. In short, LibreLog addresses privacy and cost concerns of using commercial LLMs while achieving state-of-the-arts parsing efficiency and accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Improving LLM-based Log Parsing by Learning from Errors in Reasoning TracesJialai Wang, Juncheng Lu, Jie Yang, Junjie Wang 等ASE 2025 · 被引用 2 次
- InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix CachingYilun Wang, Pengfei Chen, Haiyu Huang, Zilong He 等ICSE 2026 · 被引用 1 次
- MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context LearningJianbo Yu, Yixuan Li, Hai Xu, Kang Xu 等AAAI 2026
- VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingJinrui Sun, Tong Jia, Minghua He, Ying LiWWW 2026
- Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMsHe Yang Yuan, Xin Wang, Kundi Yao, An Ran Chen 等FSE 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- UniParser: A Unified Log Parser for Heterogeneous Log DataYudong Liu, Xu Zhang, Shilin He, Hongyu Zhang 等WWW 2022 · 被引用 148 次
- Log Parsing with Prompt-based Few-shot LearningVan-Hoang Le, Hongyu ZhangICSE 2023 · 被引用 98 次
- Guidelines for Assessing the Accuracy of Log Message Template Identification TechniquesZanis Ali Khan, Donghwan Shin, Domenico Bianculli, Lionel C. BriandICSE 2022 · 被引用 77 次
- LLMParser: An Exploratory Study on Using Large Language Models for Log ParsingZeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen 等ICSE 2024 · 被引用 72 次
相关 Paper
- SLGParser: Practical and Efficient Label-Free Log Parsing Using Large Language ModelsYibing Hu, Cong Wang, Lixin Zhao, Aimin YuICDE 2026
- Demonstration-Free: Towards More Practical Log Parsing with Large Language ModelsYi Xiao, Van-Hoang Le, Hongyu ZhangASE 2024 · 被引用 9 次
- LogParser-LLM: Advancing Efficient Log Parsing with Large Language ModelsAoxiao Zhong, Dengyao Mo, Guiyang Liu, Jinbu Liu 等KDD 2024 · 被引用 41 次
- Unleashing the True Potential of Semantic-Based Log Parsing with Pre-Trained Language ModelsVan-Hoang Le, Yi Xiao, Hongyu ZhangICSE 2025 · 被引用 6 次
- No More Labelled Examples? An Unsupervised Log Parser with LLMsJunjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. LyuFSE 2025 · 被引用 13 次
