Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug Detection
Xuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu, Shuning Liu, Xin Peng
Abstract
Regression bugs are common in real-world software projects. Although several studies have characterized various aspects of these bugs, a detailed analysis of the characteristics of code changes that introduce regression bugs is still lacking. It is also not clear whether and how regression bugs can be identified, in the era of wide usage of large language models (LLMs), in code commits during software development in the first place. To fill this gap, our study systematically analyzes 280 regression bugs from open-source Java projects, examining both the root causes and the associated development activities of regression bug-inducing changes, while also evaluating the effectiveness of LLMs in detection such commits and explaining their underlying causes.
Our findings indicate that regression bugs are usually triggered by a single atomic development activity, with feature changes or additions, and bug fixes appearing more frequently in regression bug-inducing commits. Additionally, performance improvements and refactoring are often responsible for the introduction of bugs. We develop a taxonomy that categorizes the root causes of regression bugs into two types: intrinsic errors, such as logic errors and unchecked null references, and compatibility errors, where otherwise correct changes inadvertently violate assumptions or dependencies in other parts of the system.
Furthermore, we verify that LLMs have limited ability to detect regression bugs. Based on our findings, we construct LLM4Reg that substantially improves the precision, recall, and explanation capabilities of LLM-based regression bug detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efeaca33-5197-4477-a4d2-492dcdd7b253Builds on14
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
- Enhancing Static Analysis for Practical Bug Detection: An LLM-Integrated ApproachHaonan Li, Yu Hao, Yizhuo Zhai, Zhiyun QianOOPSLA 2024 · 142 citations
- Deep just-in-time defect prediction: how far are we?Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming ZhangISSTA 2021 · 97 citations
- Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationAntonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono et al.ICSE 2020 · 81 citations
Related papers
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 1 citation
- LLM4JMH: Studying the Use of LLMs for Generating Java Performance MicrobenchmarksZongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang et al.ICSE 2026
- LLM4SZZ: Enhancing SZZ Algorithm with Context-Enhanced Assessment on Large Language ModelsLingxiao Tang, Jiakun Liu, Zhongxin Liu, Xiaohu Yang et al.ISSTA 2025 · 3 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Evaluating LLM-Based Regression Test GenerationJing Liu, Seongmin Lee, Eleonora Losiouk, Marcel BöhmeFSE 2026 · 1 citation
