Explaining mispredictions of machine learning models using rule induction
Jürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali, Satish Chandra
Abstract
While machine learning (ML) models play an increasingly prevalent role in many software engineering tasks, their prediction accuracy is often problematic. When these models do mispredict, it can be very difficult to isolate the cause. In this paper, we propose a technique that aims to facilitate the debugging process of trained statistical models. Given an ML model and a labeled data set, our method produces an interpretable characterization of the data on which the model performs particularly poorly. The output of our technique can be useful for understanding limitations of the training data or the model itself; it can also be useful for ensembling if there are multiple models with different strengths. We evaluate our approach through case studies and illustrate how it can be used to improve the accuracy of predictive models used for software engineering tasks within Facebook. We also compare our algorithm against related rule induction techniques to illustrate its advantages in the context of explaining mispredictions of machine learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41ef1d68-a06b-46c8-8441-ec690fc694b9Cited by top-tier papers6
- Graph Neural Networks for Vulnerability Detection: A Counterfactual ExplanationZhaoyang Chu, Yao Wan, Qian Li, Yang Wu et al.ISSTA 2024 · 19 citations
- Leveraging Feature Bias for Scalable Misprediction Explanation of Machine Learning ModelsJiri Gesi, Xinyun Shen, Yunfan Geng, Qihong Chen et al.ICSE 2023 · 8 citations
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 3 citations
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 3 citations
- DeciX: Explain Deep Learning Based Code Generation ApplicationsSimin Chen, Zexin Li, Wei Yang, Cong LiuFSE 2024 · 1 citation
Builds on4
- LambdaNet: Probabilistic Type Inference using Graph Neural NetworksJiayi Wei, Maruth Goyal, Greg Durrett, Isil DilligICLR 2020 · 119 citations
- TypeWriter: neural type prediction with search-based validationMichael Pradel, Georgios Gousios, Jason Liu, Satish ChandraFSE 2020 · 102 citations
- Program Synthesis Using Deduction-Guided Reinforcement LearningYanju Chen, Chenglong Wang, Osbert Bastani, Isil Dillig et al.CAV 2020 · 30 citations
- DENAS: automated rule generation by knowledge extraction from neural networksSimin Chen, Soroush Bateni, Sampath Grandhi, Xiaodi Li et al.FSE 2020 · 22 citations
Related papers
- Understanding Failures of Deep Networks via Robust Feature ExtractionSahil Singla, Besmira Nushi, Shital Shah, Ece Kamar et al.CVPR 2021
- Robust Learning of Deep Predictive Models from Noisy and Imbalanced Software Engineering DatasetsZhong Li, Minxue Pan, Yu Pei, Tian Zhang et al.ASE 2022 · 10 citations
- Efficient Understanding of Machine Learning Model MispredictionsMartin Eberlein, Jürgen Cito, Lars GrunskeASE 2025
- Pitfalls in Experiments with DNN4SE: An Analysis of the State of the PracticeSira Vegas, Sebastian G. ElbaumFSE 2023 · 4 citations
- Angler: Helping Machine Translation Practitioners Prioritize Model ImprovementsSamantha Robertson, Zijie J. Wang, Dominik Moritz, Mary Beth Kery et al.CHI 2023 · 20 citations
