Debugging Tests for Model Explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi, Been Kim
Abstract
We investigate whether post-hoc model explanations are effective for diagnosing model errors--model debugging. In response to the challenge of explaining a model's prediction, a vast array of explanation methods have been proposed. Despite increasing use, it is unclear if they are effective. To start, we categorize bugs, based on their source, into: data, model, and test-time contamination bugs. For several explanation methods, we assess their ability to: detect spurious correlation artifacts (data contamination), diagnose mislabeled training examples (data contamination), differentiate between a (partially) re-initialized model and a trained one (model contamination), and detect out-of-distribution inputs (test-time contamination). We find that the methods tested are able to diagnose a spurious background bug, but not conclusively identify mislabeled training examples. In addition, a class of methods, that modify the back-propagation algorithm are invariant to the higher layer parameters of a deep network; hence, ineffective for diagnosing model contamination. We complement our analysis with a human subject study, and find that subjects fail to identify defective models using attributions, but instead rely, primarily, on model predictions. Taken together, our results provide guidance for practitioners and researchers turning to explanations as tools for model debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ffbab03-0851-42fe-a4ca-d55c3e776a25Cited by top-tier papers46
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 167 citations
- A Consistent and Efficient Evaluation Strategy for Attribution MethodsYao Rong, Tobias Leemann, Vadim Borisov, Gjergji Kasneci et al.ICML 2022 · 138 citations
- Improving Deep Learning Interpretability by Saliency Guided TrainingAya Abdelsalam Ismail, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2021 · 121 citations
- Tracr: Compiled Transformers as a Laboratory for InterpretabilityDavid Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz et al.NeurIPS 2023 · 113 citations
- Post hoc Explanations may be Ineffective for Detecting Unknown Spurious CorrelationJulius Adebayo, Michael Muelly, Harold Abelson, Been KimICLR 2022 · 102 citations
Builds on3
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
Related papers
- How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang et al.CHI 2022 · 22 citations
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Bayes-TrEx: a Bayesian Sampling Approach to Model Transparency by ExampleSerena Booth, Yilun Zhou, Ankit Shah, Julie ShahAAAI 2021 · 20 citations
- Red Teaming Deep Neural Networks with Feature Synthesis ToolsStephen Casper, Tong Bu, Yuxiao Li, Jiawei Li et al.NeurIPS 2023 · 23 citations
- How to Probe: Simple Yet Effective Techniques for Improving Post-hoc ExplanationsSiddhartha Gairola, Moritz Böhle, Francesco Locatello, Bernt SchieleICLR 2025
