ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?
Doha Nam, Taehyoun Kim, Duksan Ryu, Jongmoon Baik
Abstract
Just-in-Time software defect prediction (JIT-SDP) plays a critical role in prioritizing risky code changes during code review and continuous integration. However, existing datasets often suffer from noisy labels and low precision in identifying bug-inducing commits. To address this, we present ReDef (Revert-based Defect dataset), a high-confidence benchmark of function-level modifications curated from 22 large-scale C/C++ projects. Defective cases are anchored by revert commits, while clean cases are validated through post-hoc history checks. Ambiguous instances are conservatively filtered out via a GPT-assisted triage process involving multiple votes and audits. This pipeline yields 3,164 defective and 10,268 clean modifications, offering substantially more reliable labels than prior resources. Beyond dataset construction, we provide a systematic evaluation of how Code Language Models (CLMs)-specifically CodeBERT, CodeT5+, UniXcoder, and Qwen2.5-reason about code modifications. We first investigate which input encodings most effectively expose change information under five different strategies. We then design four counterfactual perturbation strategies (e.g., swapping added/deleted blocks, inverting diff polarity) to serve as diagnostic probes. We posit that if models genuinely capture change semantics, such distortions should lead to a clear decline in predictive performance. Our results show that compact diff-style encodings consistently outperform whole-function formats across all CLMs, supported by rigorous statistical confirmation. However, under counterfactual tests, performance remains effectively stable, revealing that what appears to be robustness in fact reflects a reliance on superficial cues rather than true semantic understanding. These findings indicate that, at least in code-change understanding tasks, current CLMs remain limited in their ability to genuinely comprehend the relational dynamics of code modifications.
CCS Concepts: • Software and its engineering → Software defect analysis; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7620258b-013a-43c2-aa5d-06f63583fd22Builds on10
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui et al.EMNLP 2023 · 339 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
- The best of both worlds: integrating semantic features with expert features for defect prediction and localizationChao Ni, Wei Wang, Kaiwen Yang, Xin Xia et al.FSE 2022 · 76 citations
- An Empirical Comparison of Pre-Trained Models of Source CodeChangan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen et al.ICSE 2023 · 71 citations
Related papers
- SemGuard: Real-Time Semantic Evaluator for Correcting LLM-Generated CodeQinglin Wang, Zhihong Sun, Ruyun Wang, Tao Huang et al.ASE 2025 · 1 citation
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ BugsJian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu et al.ASE 2025
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng et al.FSE 2022 · 148 citations
