Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
Zhenhao Zhou, Zhuochen Huang, Yike He, Chong Wang, Jiajun Wang, Yijian Wu, Xin Peng, Yiling Lou
Abstract
The Linux kernel is a critical system, serving as the foundation for numerous systems. Bugs in the Linux kernel can cause serious consequences, affecting billions of users. Fault localization (FL), which aims at identifying the buggy code elements in software, plays an essential role in software quality assurance. While recent LLM agents have achieved promising accuracy in FL on recent benchmarks like SWE-bench, it remains unclear how well these methods perform in the Linux kernel, where FL is much more challenging due to the large-scale code base, limited observability, and diverse impact factors. In this paper, we introduce LinuxFLBench, a FL benchmark constructed from real-world Linux kernel bugs. We conduct an empirical study to assess the performance of state-of-the-art LLM agents on the Linux kernel. Our initial results reveal that existing agents struggle with this task, achieving a best top-1 accuracy of only 41.6% at file level. To address this challenge, we propose LinuxFL, an enhancement framework designed to improve FL effectiveness of LLM agents for the Linux kernel. LinuxFL substantially improves the FL accuracy of all studied agents (e.g., 7.2% - 11.2% accuracy increase) with minimal costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1db8afdb-f0a8-4a34-b15a-73fbd7f80430Cited by top-tier papers1
Ask how each one uses itBuilds on13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Boosting coverage-based fault localization via graph-based representation learningYiling Lou, Qihao Zhu, Jinhao Dong, Xia Li et al.FSE 2021 · 157 citations
- AutoCodeRover: Autonomous Program ImprovementYuntong Zhang, Haifeng Ruan, Zhiyu Fan, Abhik RoychoudhuryISSTA 2024 · 96 citations
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault LocalizationSungmin Kang, Gabin An, Shin YooFSE 2024 · 69 citations
Related papers
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
- Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for AllChenxi Huang, Alex Mathai, Feiyang Yu, Aleksandr Nogikh et al.ICML 2026
- Let the Code Speak: Incorporating Program Dynamic State for Better Method-Level Fault LocalizationYihao Qin, Shangwen Wang, Bo Lin, Xin Peng et al.ASE 2025
- OrcaLoca: An LLM Agent Framework for Software Issue LocalizationZhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang et al.ICML 2025
- Characterizing and Mitigating False-Positive Bug Reports in the Linux KernelJiashuo Tian, Dong Wang, Chen Yang, Haichi Wang et al.FSE 2026
