Small Is Beautiful: A Practical and Efficient Log Parsing Framework
Minxing Wang, Yintong Huo
摘要
Log parsing serves as the fundamental step in log analysis, splitting logs into constant templates and dynamic variables. While recent semantic-based parsers leveraging LLM have shown superior generalizability over prior syntax-based methods, their effectiveness is critically dependent on the scale of the underlying model. This dependency results in a significant performance collapse when using smaller, more practical LLMs, thereby creating a major barrier to real-world adoption where data privacy and computational constraints necessitate the use of succinct and resource-efficient models.
In a typical semantic parsing pipeline, the parsing cache is a critical component that stores the set of observed templates to quickly route incoming logs. The design of this cache is therefore paramount to the parser's overall effectiveness. Motivated by such, we improve parsing accuracy from two insights: 1) designing a more flexible cache updating strategy that can rectify prior errors, and 2) including an explicit validation process to proofread templates before they are added to the cache, preventing error injection. In particular, we propose EFParser, an unsupervised LLM-based log parser, including template extraction, template correction, and validated templates. To mitigate the impact of degraded capabilities in smaller LLMs, we designed a dual cache with an adaptive updating mechanism. When the LLM generates a new template, this module determines if it is a novel pattern or a variation of an existing one. If it's a variation, it merges the templates, thereby maintaining consistency and correcting the cache. Furthermore, we integrate a correction module that acts as a gatekeeper, validating and refining every LLM-generated template to ensure only high-quality, accurate patterns are cached. Evaluation on public large-scale datasets demonstrates that EFParser outperforms all baseline methods by an average of 12.5% across all evaluation metrics when running on smaller LLMs, with performance that even exceeds some baseline methods using large-scale LLMs, highlighting the advantages of systematic architectural design. Moreover, despite the additional processing procedures, the average processing time remains shorter than most semantic-based baselines. The superior performance on smaller LLMs combined with computational efficiency demonstrates that EFParser has significant potential for real-world deployment.
CCS Concepts: • Software and its engineering → Software creation and management.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Log-based Anomaly Detection Without Log ParsingVan-Hoang Le, Hongyu ZhangASE 2021 · 被引用 249 次
- DeepTraLog: Trace-Log Combined Microservice Anomaly Detection through Graph-based Deep LearningChenxi Zhang, Xin Peng, Chaofeng Sha, Ke Zhang 等ICSE 2022 · 被引用 163 次
相关 Paper
- SLGParser: Practical and Efficient Label-Free Log Parsing Using Large Language ModelsYibing Hu, Cong Wang, Lixin Zhao, Aimin YuICDE 2026
- EPAS: Efficient Online Log Parsing via Asynchronous Scheduling of LLM QueriesXiaolei Chen, Jia Chen, Jie Shi, Peng Wang 等ICDE 2025 · 被引用 3 次
- No More Labelled Examples? An Unsupervised Log Parser with LLMsJunjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. LyuFSE 2025 · 被引用 13 次
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language ModelsZeyang Ma, Dong Jae Kim, Tse-Hsun Peter ChenICSE 2025 · 被引用 7 次
- VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingJinrui Sun, Tong Jia, Minghua He, Ying LiWWW 2026
