The best of both worlds: integrating semantic features with expert features for defect prediction and localization
Chao Ni, Wei Wang, Kaiwen Yang, Xin Xia, Kui Liu, David Lo
Abstract
To improve software quality, just-in-time defect prediction (JIT-DP) (identifying defect-inducing commits) and just-in-time defect localization (JIT-DL) (identifying defect-inducing code lines in commits) have been widely studied by learning semantic features or expert features respectively, and indeed achieved promising performance. Semantic features and expert features describe code change commits from different aspects, however, the best of the two features have not been fully explored together to boost the just-in-time defect prediction and localization in the literature yet. Additional, JIT-DP identifies defects at the coarse commit level, while as the consequent task of JIT-DP, JIT-DL cannot achieve the accurate localization of defect-inducing code lines in a commit without JIT-DP. We hypothesize that the two JIT tasks can be combined together to boost the accurate prediction and localization of defect-inducing commits by integrating semantic features with expert features. Therefore, we propose to build a unified model, JIT-Fine, for the just-in-time defect prediction and localization by leveraging the best of semantic features and expert features. To assess the feasibility of JIT-Fine, we first build a large-scale line-level manually labeled dataset, JIT-Defects4J. Then, we make a comprehensive comparison with six state-of-the-art baselines under various settings using ten performance measures grouped into two types: effort-agnostic and effort-aware. The experimental results indicate that JIT-Fine can outperform all state-of-the-art baselines on both JIT-DP and JITDL tasks in terms of ten performance measures with a substantial improvement (i.e., 10%-629% in terms of effort-agnostic measures on JIT-DP, 5%-54% in terms of effort-aware measures on JIT-DP, and 4%-117% in terms of effort-aware measures on JIT-DL).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcf7e1d7-8e44-47da-9af2-d0920b4785b8Cited by top-tier papers10
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao et al.FSE 2023 · 42 citations
- CREF: An LLM-Based Conversational Software Repair Framework for Programming TutorsBoyang Yang, Haoye Tian, Weiguo Pian, Haoran Yu et al.ISSTA 2024 · 26 citations
- Pre-training Code Representation with Semantic Flow Graph for Effective Bug LocalizationYali Du, Zhongxing YuFSE 2023 · 18 citations
- LogSD: Detecting Anomalies from System Logs through Self-Supervised Learning and Frequency-Based MaskingYongzheng Xie, Hongyu Zhang, Muhammad Ali BabarFSE 2024 · 14 citations
Builds on6
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- Deep just-in-time defect prediction: how far are we?Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming ZhangISSTA 2021 · 97 citations
- An investigation of cross-project learning in online just-in-time software defect predictionSadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral et al.ICSE 2020 · 49 citations
- Evaluating SZZ Implementations Through a Developer-informed OracleGiovanni Rosa, Luca Pascarella, Simone Scalabrino, Rosalia Tufano et al.ICSE 2021 · 42 citations
- Automating the removal of obsolete TODO commentsZhipeng Gao, Xin Xia, David Lo, John C. Grundy et al.FSE 2021 · 34 citations
Related papers
- JIT-Smart: A Multi-task Learning Framework for Just-in-Time Defect Prediction and LocalizationXiangping Chen, Furen Xu, Yuan Huang, Neng Zhang et al.FSE 2024 · 11 citations
- A Practical Human Labeling Method for Online Just-in-Time Software Defect PredictionLiyan Song, Leandro L. Minku, Cong Teng, Xin YaoFSE 2023 · 6 citations
- ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?Doha Nam, Taehyoun Kim, Duksan Ryu, Jongmoon BaikFSE 2026
- NeuroJIT: Improving Just-In-Time Defect Prediction Using Neurophysiological and Empirical Perceptions of Modern DevelopersGichan Lee, Hansae Ju, Scott Uk-Jin LeeASE 2024
- PyExplainer: Explaining the Predictions of Just-In-Time Defect ModelsChanathip Pornprasit, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Michael Fu et al.ASE 2021 · 52 citations
