Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
Xin Wang, Zhenhao Li, Zishuo Ding
摘要
Logging code is written by developers to capture system runtime behavior and plays a vital role in debugging, performance analysis, and system monitoring. However, defects in logging code can undermine the usefulness of logs and lead to misinterpretations. Although prior work has identified several logging defect patterns and provided valuable insights into logging practices, these studies often focus on a narrow range of defect patterns derived from limited sources (e.g., commit histories) and lack a systematic and comprehensive analysis. Moreover, large language models (LLMs) have demonstrated promising generalization and reasoning capabilities across a variety of code-related tasks, yet their potential for detecting logging code defects remains largely unexploredIn this paper, we derive a comprehensive taxonomy of logging code defects, which encompasses seven logging code defect patterns with 14 detailed scenarios. We further construct a benchmark dataset, Defects4Log, consisting of 164 developer-verified real-world logging defects. Then we propose an automated framework that leverages various prompting strategies and contextual information to evaluate LLMs’ capability in detecting and reasoning logging code defects. Experimental results reveal that LLMs generally struggle to accurately detect and reason logging code defects based on the source code only. However, incorporating proper knowledge (e.g., detailed scenarios of defect patterns) can lead to 10.9% improvement in detection accuracy. Overall, our findings provide actionable guidance for practitioners to avoid common defect patterns and establish a foundation for improving LLM-based reasoning in logging code defect detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance ModelingXin Wang, Zhenhao Li, Zishuo DingICSE 2026
- Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMsHe Yang Yuan, Xin Wang, Kundi Yao, An Ran Chen 等FSE 2026
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- LLMParser: An Exploratory Study on Using Large Language Models for Log ParsingZeyang Ma, An Ran Chen, Dong Jae Kim, Tse-Hsun Chen 等ICSE 2024 · 被引用 72 次
- DeepLV: Suggesting Log Levels Using Ordinal Based Neural NetworksZhenhao Li, Heng Li, Tse-Hsun Peter Chen, Weiyi ShangICSE 2021 · 被引用 44 次
- Where Shall We Log? Studying and Suggesting Logging Locations in Code BlocksZhenhao Li, Tse-Hsun Chen, Weiyi ShangASE 2020 · 被引用 41 次
相关 Paper
- UniLog: Automatic Logging via LLM and In-Context LearningJunjielong Xu, Ziang Cui, Yuan Zhao, Xu Zhang 等ICSE 2024 · 被引用 56 次
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 被引用 2 次
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- LILAC: Log Parsing using LLMs with Adaptive Parsing CacheZhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li 等FSE 2024 · 被引用 85 次
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement FrameworkLuping Zhang, Chao Chen, Dapeng Yan, Hui Xu 等FSE 2026
