Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
Xin Wang, Zhenhao Li, Zishuo Ding
Abstract
Logging code is written by developers to capture system runtime behavior and plays a vital role in debugging, performance analysis, and system monitoring. However, defects in logging code can undermine the usefulness of logs and lead to misinterpretations. Although prior work has identified several logging defect patterns and provided valuable insights into logging practices, these studies often focus on a narrow range of defect patterns derived from limited sources (e.g., commit histories) and lack a systematic and comprehensive analysis. Moreover, large language models (LLMs) have demonstrated promising generalization and reasoning capabilities across a variety of code-related tasks, yet their potential for detecting logging code defects remains largely unexploredIn this paper, we derive a comprehensive taxonomy of logging code defects, which encompasses seven logging code defect patterns with 14 detailed scenarios. We further construct a benchmark dataset, Defects4Log, consisting of 164 developer-verified real-world logging defects. Then we propose an automated framework that leverages various prompting strategies and contextual information to evaluate LLMs’ capability in detecting and reasoning logging code defects. Experimental results reveal that LLMs generally struggle to accurately detect and reason logging code defects based on the source code only. However, incorporating proper knowledge (e.g., detailed scenarios of defect patterns) can lead to 10.9% improvement in detection accuracy. Overall, our findings provide actionable guidance for practitioners to avoid common defect patterns and establish a foundation for improving LLM-based reasoning in logging code defect detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5985447c-8770-4116-a72e-1f5c8d9198e2Cited by top-tier papers2
- LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance ModelingXin Wang, Zhenhao Li, Zishuo DingICSE 2026
- Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMsHe Yang Yuan, Xin Wang, Kundi Yao, An Ran Chen et al.FSE 2026
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- LLMParser: An Exploratory Study on Using Large Language Models for Log ParsingZeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen et al.ICSE 2024 · 72 citations
- DeepLV: Suggesting Log Levels Using Ordinal Based Neural NetworksZhenhao Li, Heng Li, Tse-Hsun Peter Chen, Weiyi ShangICSE 2021 · 44 citations
- Where Shall We Log? Studying and Suggesting Logging Locations in Code BlocksZhenhao Li, Tse-Hsun Chen, Weiyi ShangASE 2020 · 41 citations
Related papers
- UniLog: Automatic Logging via LLM and In-Context LearningJunjielong Xu, Ziang Cui, Yuan Zhao, Xu Zhang et al.ICSE 2024 · 56 citations
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 2 citations
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- LILAC: Log Parsing using LLMs with Adaptive Parsing CacheZhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li et al.FSE 2024 · 85 citations
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement FrameworkLuping Zhang, Chao Chen, Dapeng Yan, Hui Xu et al.FSE 2026
