Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model Explanation
Thien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun Sakuma
Abstract
We aim to explain a black-box classifier with the form: "data X is classified as class Y because X has A, B and does not have C" in which A, B, and C are high-level concepts. The challenge is that we have to discover in an unsupervised manner a set of concepts, i.e., A, B and C, that is useful for explaining the classifier. We first introduce a structural generative model that is suitable to express and discover such concepts. We then propose a learning process that simultaneously learns the data distribution and encourages certain concepts to have a large causal influence on the classifier output. Our method also allows easy integration of user's prior knowledge to induce high interpretability of concepts. Finally, using multiple datasets, we demonstrate that the proposed method can discover useful concepts for explanation in this form.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9de39d05-1f4e-492c-ba44-a43e2dd86866Builds on4
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
- Benchmarks, Algorithms, and Metrics for Hierarchical DisentanglementAndrew Slavin Ross, Finale Doshi-VelezICML 2021 · 15 citations
- PatchVAE: Learning Local Latent Codes for RecognitionKamal Gupta, Saurabh Singh, Abhinav ShrivastavaCVPR 2020
Related papers
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf et al.NeurIPS 2024 · 37 citations
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng et al.CVPR 2021
- Debiasing Concept-based Explanations with Causal AnalysisMohammad Taha Bahadori, David HeckermanICLR 2021 · 8 citations
- Learning Discrete Concepts in Latent Hierarchical ModelsLingjing Kong, Guangyi Chen, Biwei Huang, Eric P. Xing et al.NeurIPS 2024 · 20 citations
