LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation
Fei Teng, Haoyang Li, Lei Chen
Abstract
Modern computing systems, such as HDFS and Spark, produce vast quantities of logs that developers use for tasks like anomaly detection and error analysis. To simplify log analysis, template generation methods have been proposed to standardize log formats, transforming unstructured data into structured templates. Existing heuristic-based methods and neural network-based methods suffer from low accuracy problems due to the reliance on handcrafted heuristics or specific log patterns in training sets. Recently, large language models (LLMs) have shown great potential in log template generation. However, they often struggle with ambiguous, complex, or highly specific log content, which can lead to errors in generating accurate templates. To address these challenges, we propose LLMLog, a multi-round annotation framework with adaptive in-context learning. We first propose an edit-distance-based similarity metric to evaluate log similarity. Then, we introduce a method to select the most informative k unlabeled logs for annotation by considering both the representativeness of the logs and the confidence of LLM predictions. Additionally, we design an adaptive context selection strategy that adaptively selects labeled logs to ensure comprehensive keyword coverage for unlabeled logs. These labeled logs serve as the context for LLMs to better understand the unlabeled logs, thereby enhancing the accuracy of template generation. Extensive experiments on sixteen datasets demonstrate that LLMLog outperforms the state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a65af3f-02dd-4f6c-b36f-b9778de89f66Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 1,823 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- UniParser: A Unified Log Parser for Heterogeneous Log DataYudong Liu, Xu Zhang, Shilin He, Hongyu Zhang et al.WWW 2022 · 148 citations
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan et al.VLDB 2023 · 127 citations
Related papers
- Semantic Curriculum for Anomaly Detection: A Unified Language-Driven Meta-Optimization FrameworkKai Tan, Yangliu Du, Dongyang Zhan, Haining Yu et al.INFOCOM 2026
- LILAC: Log Parsing using LLMs with Adaptive Parsing CacheZhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li et al.FSE 2024 · 85 citations
- DivLog: Log Parsing with Prompt Enhanced In-Context LearningJunjielong Xu, Ruichun Yang, Yintong Huo, Chengyu Zhang et al.ICSE 2024 · 54 citations
- Efficient Zero-Shot and Label-free Log Anomaly Detection for Resource-Constrained SystemsZuohan Wu, Jiachuan Wang, Libin Zheng, Yongqi Zhang et al.ICDE 2026
- LLMParser: An Exploratory Study on Using Large Language Models for Log ParsingZeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen et al.ICSE 2024 · 72 citations
