Deep just-in-time defect prediction: how far are we?
Zhengran Zeng, Yuqun Zhang, Haotian Zhang, Lingming Zhang
摘要
Defect prediction aims to automatically identify potential defective code with minimal human intervention and has been widely studied in the literature. Just-in-Time (JIT) defect prediction focuses on program changes rather than whole programs, and has been widely adopted in continuous testing. CC2Vec, state-of-the-art JIT defect prediction tool, first constructs a hierarchical attention network (HAN) to learn distributed vector representations of both code additions and deletions, and then concatenates them with two other embedding vectors representing commit messages and overall code changes extracted by the existing DeepJIT approach to train a model for predicting whether a given commit is defective. Although CC2Vec has been shown to be the state of the art for JIT defect prediction, it was only evaluated on a limited dataset and not compared with all representative baselines. Therefore, to further investigate the efficacy and limitations of CC2Vec, this paper performs an extensive study of CC2Vec on a large-scale dataset with over 310,370 changes (8.3 X larger than the original CC2Vec dataset). More specifically, we also empirically compare CC2Vec against DeepJIT and representative traditional JIT defect prediction techniques. The experimental results show that CC2Vec cannot consistently outperform DeepJIT, and neither of them can consistently outperform traditional JIT defect prediction. We also investigate the impact of individual traditional defect prediction features and find that the added-line-number feature outperforms other traditional features. Inspired by this finding, we construct a simplistic JIT defect prediction approach which simply adopts the added-line-number feature with the logistic regression classifier. Surprisingly, such a simplistic approach can outperform CC2Vec and DeepJIT in defect prediction, and can be 81k X/120k X faster in training/testing. Furthermore, the paper also provides various practical guidelines for advancing JIT defect prediction in the near future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li 等ISSTA 2022 · 被引用 142 次
- Free Lunch for Testing: Fuzzing Deep-Learning Libraries from Open SourceAnjiang Wei, Yinlin Deng, Chenyuan Yang, Lingming ZhangICSE 2022 · 被引用 91 次
- Fuzzing deep-learning libraries via automated relational API inferenceYinlin Deng, Chenyuan Yang, Anjiang Wei, Lingming ZhangFSE 2022 · 被引用 83 次
- The best of both worlds: integrating semantic features with expert features for defect prediction and localizationChao Ni, Wei Wang, Kaiwen Yang, Xin Xia 等FSE 2022 · 被引用 76 次
- Fast Changeset-based Bug Localization with BERTAgnieszka Ciborowska, Kostadin DamevskiICSE 2022 · 被引用 53 次
它引用的顶会 Paper3
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier 等ASE 2020 · 被引用 91 次
- An investigation of cross-project learning in online just-in-time software defect predictionSadia Tabassum, Leandro L. Minku, Danyi Feng, George G. Cabral 等ICSE 2020 · 被引用 49 次
相关 Paper
- NeuroJIT: Improving Just-In-Time Defect Prediction Using Neurophysiological and Empirical Perceptions of Modern DevelopersGichan Lee, Hansae Ju, Scott Uk-Jin LeeASE 2024
- CCRep: Learning Code Change Representations via Pre-Trained Code Model and Query BackZhongxin Liu, Zhijie Tang, Xin Xia, Xiaohu YangICSE 2023 · 被引用 25 次
- JIT-Smart: A Multi-task Learning Framework for Just-in-Time Defect Prediction and LocalizationXiangping Chen, Furen Xu, Yuan Huang, Neng Zhang 等FSE 2024 · 被引用 11 次
- DeepCVA: Automated Commit-level Vulnerability Assessment with Deep Multi-task LearningTriet Huynh Minh Le, David Hin, Roland Croft, Muhammad Ali BabarASE 2021 · 被引用 62 次
- VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit IdentificationXing Cui, Jingzheng Wu, Wenxiang Ou, Tianyue Luo 等AAAI 2026
