Small Is Beautiful: A Practical and Efficient Log Parsing Framework
Minxing Wang, Yintong Huo
Abstract
Log parsing serves as the fundamental step in log analysis, splitting logs into constant templates and dynamic variables. While recent semantic-based parsers leveraging LLM have shown superior generalizability over prior syntax-based methods, their effectiveness is critically dependent on the scale of the underlying model. This dependency results in a significant performance collapse when using smaller, more practical LLMs, thereby creating a major barrier to real-world adoption where data privacy and computational constraints necessitate the use of succinct and resource-efficient models.
In a typical semantic parsing pipeline, the parsing cache is a critical component that stores the set of observed templates to quickly route incoming logs. The design of this cache is therefore paramount to the parser's overall effectiveness. Motivated by such, we improve parsing accuracy from two insights: 1) designing a more flexible cache updating strategy that can rectify prior errors, and 2) including an explicit validation process to proofread templates before they are added to the cache, preventing error injection. In particular, we propose EFParser, an unsupervised LLM-based log parser, including template extraction, template correction, and validated templates. To mitigate the impact of degraded capabilities in smaller LLMs, we designed a dual cache with an adaptive updating mechanism. When the LLM generates a new template, this module determines if it is a novel pattern or a variation of an existing one. If it's a variation, it merges the templates, thereby maintaining consistency and correcting the cache. Furthermore, we integrate a correction module that acts as a gatekeeper, validating and refining every LLM-generated template to ensure only high-quality, accurate patterns are cached. Evaluation on public large-scale datasets demonstrates that EFParser outperforms all baseline methods by an average of 12.5% across all evaluation metrics when running on smaller LLMs, with performance that even exceeds some baseline methods using large-scale LLMs, highlighting the advantages of systematic architectural design. Moreover, despite the additional processing procedures, the average processing time remains shorter than most semantic-based baselines. The superior performance on smaller LLMs combined with computational efficiency demonstrates that EFParser has significant potential for real-world deployment.
CCS Concepts: • Software and its engineering → Software creation and management.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ec670a1-fe5b-4e7f-8a86-fbf1209103a2Builds on16
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 1,823 citations
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Log-based Anomaly Detection Without Log ParsingVan-Hoang Le, Hongyu ZhangASE 2021 · 249 citations
- DeepTraLog: Trace-Log Combined Microservice Anomaly Detection through Graph-based Deep LearningChenxi Zhang, Xin Peng, Chaofeng Sha, Ke Zhang et al.ICSE 2022 · 163 citations
Related papers
- SLGParser: Practical and Efficient Label-Free Log Parsing Using Large Language ModelsYibing Hu, Cong Wang, Lixin Zhao, Aimin YuICDE 2026
- EPAS: Efficient Online Log Parsing via Asynchronous Scheduling of LLM QueriesXiaolei Chen, Jia Chen, Jie Shi, Peng Wang et al.ICDE 2025 · 3 citations
- No More Labelled Examples? An Unsupervised Log Parser with LLMsJunjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. LyuFSE 2025 · 13 citations
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language ModelsZeyang Ma, Dong Jae Kim, Tse-Hsun Peter ChenICSE 2025 · 7 citations
- VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingJinrui Sun, Tong Jia, Minghua He, Ying LiWWW 2026
