Meaningfully debugging model mistakes using conceptual counterfactual explanations
Abubakar Abid, Mert Yüksekgönül, James Zou
Abstract
Understanding and explaining the mistakes made by trained models is critical to many machine learning objectives, such as improving robustness, addressing concept drift, and mitigating biases. However, this is often an ad hoc process that involves manually looking at the model's mistakes on many test samples and guessing at the underlying reasons for those incorrect predictions. In this paper, we propose a systematic approach, conceptual counterfactual explanations (CCE), that explains why a classifier makes a mistake on a particular test sample(s) in terms of human-understandable concepts (e.g. this zebra is misclassified as a dog because of faint stripes). We base CCE on two prior ideas: counterfactual explanations and concept activation vectors, and validate our approach on wellknown pretrained models, showing that it explains the models' mistakes meaningfully. In addition, for new models trained on data with spurious correlations, CCE accurately identifies the spurious correlation as the cause of model mistakes from a single misclassified test sample. On two challenging medical applications, CCE generated useful insights, confirmed by clinicians, into biases and mistakes the model makes in real-world settings. The code for CCE is publicly available at https://github.com/ mertyg/debug-mistakes-cce .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d20b9c0-985a-4b54-b851-68fe47f4c446Cited by top-tier papers21
- Probabilistic Concept Bottleneck ModelsEunji Kim, Dahuin Jung, Sangha Park, Siwon Kim et al.ICML 2023 · 108 citations
- Discover and Cure: Concept-aware Mitigation of Spurious CorrelationShirley Wu, Mert Yüksekgönül, Linjun Zhang, James ZouICML 2023 · 97 citations
- Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?Sonia Laguna, Ricards Marcinkevics, Moritz Vandenhirtz, Julia E. VogtNeurIPS 2024 · 39 citations
- Seeing is not Believing: Robust Reinforcement Learning against Spurious CorrelationWenhao Ding, Laixi Shi, Yuejie Chi, Ding ZhaoNeurIPS 2023 · 39 citations
- Post-hoc Concept Bottleneck ModelsMert Yüksekgönül, Maggie Wang, James ZouICLR 2023 · 37 citations
Builds on2
Related papers
- Grounding Counterfactual Explanation of Image Classifiers to Textual Concept SpaceSiwon Kim, Jinoh Oh, Sungjin Lee, Seunghak Yu et al.CVPR 2023
- Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsTownim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong et al.ICCV 2025
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo et al.ICCV 2023 · 15 citations
- CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-LinesArjun R. Akula, Shuai Wang, Song-Chun ZhuAAAI 2020 · 102 citations
