Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
Jiashuo Tian, Dong Wang, Chen Yang, Haichi Wang, Zan Wang, Junjie Chen
Abstract
False-positive bug reports represent a significant yet underexplored challenge in the development and maintenance of the Linux kernel. They occur when correct system behavior is mistakenly flagged as a defect, consuming developer effort without leading to actual code improvements. Such reports can mislead developers, waste debugging resources, and delay the resolution of real bugs. In this paper, we present the first comprehensive empirical study of false-positive bug reports in the Linux kernel. We manually construct a dataset of 2,006 bug reports comprising 1,509 genuine bugs and 497 false positives collected from Bugzilla and Syzkaller. Our analysis indicates that false positives demand effort comparable to real bugs, often requiring extended discussions and non-trivial closure time. They occur in several components, especially File Systems and Drivers, mainly due to external dependencies and semantic misunderstandings. To address this challenge, we evaluate large language models (LLMs) for automated false-positive bug report mitigation. Among various prompting strategies, retrieval-augmented generation (RAG) performs best, achieving 91% recall and an F1 score of 88%. These findings highlight the non-negligible cost of false positive bug reports and show the promise of LLMs for more efficient false positive mitigation in the Linux kernel.
CCS Concepts: • Software and its engineering → Software libraries and repositories; Software defect analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- kAFL: Hardware-Assisted Feedback Fuzzing for OS KernelsSergej Schumilo, Cornelius Aschermann, Robert Gawlik, Sebastian Schinzel et al.USENIX Security 2017 · 324 citations
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 273 citations
- MoonShine: Optimizing OS Fuzzer Seed Selection with Trace DistillationShankara Pailoor, Andrew Aday, Suman JanaUSENIX Security 2018 · 180 citations
Related papers
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault LocalizationSungmin Kang, Gabin An, Shin YooFSE 2024 · 69 citations
- An In-depth Analysis of Duplicated Linux Kernel Bug ReportsDongliang Mu, Yuhang Wu, Yueqi Chen, Zhenpeng Lin et al.NDSS 2022
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 5 citations
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi et al.ISSTA 2025 · 53 citations
- KernelGPT: Enhanced Kernel Fuzzing via Large Language ModelsChenyuan Yang, Zijie Zhao, Lingming ZhangASPLOS 2025 · 45 citations
