Grounding Counterfactual Explanation of Image Classifiers to Textual Concept Space
Siwon Kim, Jinoh Oh, Sungjin Lee, Seunghak Yu, Jaeyoung Do, Tara Taghavi
Abstract
Concept-based explanation aims to provide concise and human-understandable explanations of an image classifier. However, existing concept-based explanation methods typically require a significant amount of manually collected concept-annotated images. This is costly and runs the risk of human biases being involved in the explanation. In this paper, we propose Counterfactual explanation with text-driven concepts (CounTEX), where the concepts are defined only from text by leveraging a pre-trained multimodal joint embedding space without additional conceptannotated datasets. A conceptual counterfactual explanation is generated with text-driven concepts. To utilize the text-driven concepts defined in the joint embedding space to interpret target classifier outcome, we present a novel projection scheme for mapping the two spaces with a simple yet effective implementation. We show that CounTEX generates faithful explanations that provide a semantic understanding of model decision rationale robust to human bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a9974de-7fe5-4110-a91d-9d2959c954b1Cited by top-tier papers9
- Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?Sonia Laguna, Ricards Marcinkevics, Moritz Vandenhirtz, Julia E. VogtNeurIPS 2024 · 39 citations
- MCPNet: An Interpretable Classifier via Multi-Level Concept PrototypesBor-Shiun Wang, Chien-Yi Wang, Wei-Chen ChiuCVPR 2024 · 11 citations
- Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation ModelsHengyi Wang, Shiwei Tan, Hao WangICML 2024 · 9 citations
- CE-FAM: Concept-Based Explanation via Fusion of Activation MapsMichihiro Kuroki, Toshihiko YamasakiICCV 2025 · 3 citations
- Causality-aligned Prompt Learning via Diffusion-based Counterfactual GenerationXinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong et al.ACM MM 2025 · 3 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano et al.SIGGRAPH 2022 · 501 citations
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald et al.ICCV 2021 · 181 citations
Related papers
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsTownim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong et al.ICCV 2025
- Zero-Shot Natural Language ExplanationsFawaz Sammani, Nikos DeligiannisICLR 2025
- Meaningfully debugging model mistakes using conceptual counterfactual explanationsAbubakar Abid, Mert Yüksekgönül, James ZouICML 2022 · 75 citations
- DISSECT: Disentangled Simultaneous Explanations via Concept TraversalsAsma Ghandeharioun, Been Kim, Chun-Liang Li, Brendan Jou et al.ICLR 2022 · 58 citations
