Deep just-in-time defect prediction: how far are we?
Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming Zhang
Abstract
Defect prediction aims to automatically identify potential defective code with minimal human intervention and has been widely studied in the literature. Just-in-Time (JIT) defect prediction focuses on program changes rather than whole programs, and has been widely adopted in continuous testing. CC2Vec, state-of-the-art JIT defect prediction tool, first constructs a hierarchical attention network (HAN) to learn distributed vector representations of both code additions and deletions, and then concatenates them with two other embedding vectors representing commit messages and overall code changes extracted by the existing DeepJIT approach to train a model for predicting whether a given commit is defective. Although CC2Vec has been shown to be the state of the art for JIT defect prediction, it was only evaluated on a limited dataset and not compared with all representative baselines. Therefore, to further investigate the efficacy and limitations of CC2Vec, this paper performs an extensive study of CC2Vec on a large-scale dataset with over 310,370 changes (8.3 X larger than the original CC2Vec dataset). More specifically, we also empirically compare CC2Vec against DeepJIT and representative traditional JIT defect prediction techniques. The experimental results show that CC2Vec cannot consistently outperform DeepJIT, and neither of them can consistently outperform traditional JIT defect prediction. We also investigate the impact of individual traditional defect prediction features and find that the added-line-number feature outperforms other traditional features. Inspired by this finding, we construct a simplistic JIT defect prediction approach which simply adopts the added-line-number feature with the logistic regression classifier. Surprisingly, such a simplistic approach can outperform CC2Vec and DeepJIT in defect prediction, and can be 81k X/120k X faster in training/testing. Furthermore, the paper also provides various practical guidelines for advancing JIT defect prediction in the near future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8852f73e-d8ef-46d6-bb75-96f33661d6c8Cited by top-tier papers13
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li et al.ISSTA 2022 · 142 citations
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 91 citations
- Fuzzing deep-learning libraries via automated relational API inferenceYinlin Deng, Chenyuan Yang, Anjiang Wei, Lingming ZhangFSE 2022 · 83 citations
- The best of both worlds: integrating semantic features with expert features for defect prediction and localizationChao Ni, Wei Wang, Kaiwen Yang, Xin Xia et al.FSE 2022 · 76 citations
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 53 citations
Builds on3
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- An investigation of cross-project learning in online just-in-time software defect predictionSadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral et al.ICSE 2020 · 49 citations
Related papers
- NeuroJIT: Improving Just-In-Time Defect Prediction Using Neurophysiological and Empirical Perceptions of Modern DevelopersGichan Lee, Hansae Ju, Scott Uk-Jin LeeASE 2024
- CCRep: Learning Code Change Representations via Pre-Trained Code Model and Query BackZhongxin Liu, Zhijie Tang, Xin Xia, Xiaohu YangICSE 2023 · 25 citations
- JIT-Smart: A Multi-task Learning Framework for Just-in-Time Defect Prediction and LocalizationXiangping Chen, Furen Xu, Yuan Huang, Neng Zhang et al.FSE 2024 · 11 citations
- DeepCVA: Automated Commit-level Vulnerability Assessment with Deep Multi-task LearningTriet Huynh Minh Le, David Hin, Roland Croft, Muhammad Ali BabarASE 2021 · 62 citations
- VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit IdentificationXing Cui, Jingzheng Wu, Wenxiang Ou, Tianyue Luo et al.AAAI 2026
