Generative causal explanations of black-box classifiers
Matthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell, Mark A. Davenport
Abstract
We develop a method for generating causal post-hoc explanations of black-box classifiers based on a learned low-dimensional representation of the data. The explanation is causal in the sense that changing learned latent factors produces a change in the classifier output statistics. To construct these explanations, we design a learning framework that leverages a generative model and information-theoretic measures of causal influence. Our objective function encourages both the generative model to faithfully represent the data distribution and the latent factors to have a large causal influence on the classifier output. Our method learns both global and local explanations, is compatible with any classifier that admits class probabilities and a gradient, and does not require labeled attributes or knowledge of causal structure. Using carefully controlled test cases, we provide intuition that illuminates the function of our causal objective. We then demonstrate the practical utility of our method on image recognition tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8ca5dff-46b7-43ee-9216-9cafa6381eceCited by top-tier papers16
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald et al.ICCV 2021 · 181 citations
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Learning to Receive Help: Intervention-Aware Concept Embedding ModelsMateo Espinosa Zarlenga, Katie Collins, Krishnamurthy Dvijotham, Adrian Weller et al.NeurIPS 2023 · 56 citations
- FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI MethodsRobin Hesse, Simone Schaub-Meyer, Stefan RothICCV 2023 · 50 citations
Builds on2
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
Related papers
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- Explanations for Occluded ImagesHana Chockler, Daniel Kroening, Youcheng SunICCV 2021 · 23 citations
- Accurate Explanation Model for Image Classifiers using Class Association EmbeddingRuitao Xie, Jingbang Chen, Limai Jiang, Rui Xiao et al.ICDE 2024 · 12 citations
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 116 citations
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
