A Knowledge Enhanced Large Language Model for Bug Localization
Yue Li, Bohan Liu, Ting Zhang, Zhiqi Wang, David Lo, Lanxin Yang, Jun Lyu, He Zhang
Abstract
A significant number of bug reports are generated every day as software systems continue to develop. Large Language Models (LLMs) have been used to correlate bug reports with source code to locate bugs automatically. The existing research has shown that LLMs are effective for bug localization and can increase software development efficiency. However, these studies still have two limitations. First, these models fail to capture context information about bug reports and source code. Second, these models are unable to understand the domain-specific expertise inherent to particular projects, such as version information in projects that are composed of alphanumeric characters without any semantic meaning. To address these challenges, we propose a K nowledge E nhanced P re- T rained model using project documents and historical code, called KEPT , for bug localization. Project documents record, revise, and restate project information that provides rich semantic information about those projects. Historical code contains rich code semantic information that can enhance the reasoning ability of LLMs. Specifically, we construct knowledge graphs from project documents and source code. Then, we introduce knowledge graphs to the LLM through soft-position embedding and visible matrices, enhancing its contextual and professional reasoning ability. To validate our model, we conducted a series of experiments on seven open-source software projects with over 6,000 bug reports. Compared with the traditional model (Locus), KEPT performs better by 33.2% to 59.5% in terms of mean reciprocal rank, mean average precision, and Top@N. Compared with the best-performing non-commercial LLM (CodeT5), KEPT achieves an improvement of 36.6% to 63.7%. Compared to the state-of-the-art commercial LLM developed by OpenAI, called text-embedding-ada-002 , KEPT achieves an average improvement of 7.8% to 17.4%. The results indicate that introducing knowledge graphs contributes to enhance the effectiveness of the LLM in bug localization.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3f1c1a01-511f-4fa3-bc8e-64d02388a9b9Cited by top-tier papers1
Ask how each one uses itRelated papers
- Location is Key: Leveraging LLM for Functional Bug Localization in Verilog DesignBingkun Yao, Ning Wang, Jie Zhou, Xi Wang et al.DAC 2025 · 9 citations
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 2 citations
- A Comparative Study of Transformer-Based Neural Text Representation Techniques on Bug TriagingAtish Kumar Dipongkor, Kevin MoranASE 2023 · 9 citations
- Debugging Engine Enhanced by Prior Knowledge: Can We Teach LLM How to Debug?Kunyi Li, Sai Wu, Xiu Tang, Chang Yao et al.FSE 2026 · 1 citation
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 53 citations
