PyExplainer: Explaining the Predictions of Just-In-Time Defect Models
Chanathip Pornprasit, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Michael Fu, Patanamon Thongtanunam
Abstract
Just-In-Time (JIT) defect prediction (i.e., an AI/ML model to predict defect-introducing commits) is proposed to help developers prioritize their limited Software Quality Assurance (SQA) resources on the most risky commits. However, the explainability of JIT defect models remains largely unexplored (i.e., practitioners still do not know why a commit is predicted as defect-introducing). Recently, LIME has been used to generate explanations for any AI/ML models. However, the random perturbation approach used by LIME to generate synthetic neighbors is still suboptimal, i.e., generating synthetic neighbors that may not be similar to an instance to be explained, producing low accuracy of the local models, leading to inaccurate explanations for just-in-time defect models. In this paper, we propose PyExplainer-i.e., a local rule-based model-agnostic technique for generating explanations (i.e., why a commit is predicted as defective) of JIT defect models. Through a case study of two open-source software projects, we find that our PyExplainer produces (1) synthetic neighbors that are 41%-45% more similar to an instance to be explained; (2) 18%-38% more accurate local models; and (3) explanations that are 69%-98% more unique and 17%-54% more consistent with the actual characteristics of defect-introducing commits in the future than LIME (a state-ofthe-art model-agnostic technique). This could help practitioners focus on the most important aspects of the commits to mitigate the risk of being defect-introducing. Thus, the contributions of this paper build an important step towards Explainable AI for Software Engineering, making software analytics more explainable and actionable. Finally, we publish our PyExplainer as a Python package to support practitioners and researchers ( https://github.com/awsm-research/PyExplainer ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbdc0df0-30e6-4d56-8d40-e214bc17cfd7Cited by top-tier papers6
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen et al.FSE 2022 · 206 citations
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo et al.ICSE 2024 · 27 citations
- Explaining Software Bugs Leveraging Code Structures in Neural Machine TranslationParvez Mahbub, Ohiduzzaman Shuvo, Mohammad Masudur RahmanICSE 2023 · 22 citations
- Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!Xu Yang, Shaowei Wang, Yi Li, Shaohua WangICSE 2023 · 19 citations
- What You See is What You Get: Attention-Based Self-Guided Automatic Unit Test GenerationXin Yin, Chao Ni, Xiaodan Xu, Xiaohu YangICSE 2025 · 8 citations
Builds on1
Related papers
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang et al.USENIX Security 2025
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 6 citations
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 56 citations
- Shahin: Faster Algorithms for Generating Explanations for Multiple PredictionsSona Hasani, Saravanan Thirumuruganathan, Nick Koudas, Gautam DasSIGMOD 2021 · 1 citation
- The best of both worlds: integrating semantic features with expert features for defect prediction and localizationChao Ni, Wei Wang, Kaiwen Yang, Xin Xia et al.FSE 2022 · 76 citations
