Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification Models
Townim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong, Minh-Son To, Anton van den Hengel, Johan W. Verjans, Zhibin Liao
Abstract
Counterfactual explanations (CFE) for deep image classifiers aim to reveal how minimal input changes lead to different model decisions, providing critical insights for model interpretation and improvement. However, existing CFE methods often rely on additional image encoders and generative models to create plausible images, neglecting the classifier's own feature space and decision boundaries. As such, they do not explain the intrinsic feature space and decision boundaries learned by the classifier. To address this limitation, we propose Mirror-CFE, a novel method that generates faithful counterfactual explanations by operating directly in the classifier's feature space, treating decision boundaries as mirrors that ``reflect'' feature representations in the mirror. Mirror-CFE learns a mapping function from feature space to image space while preserving distance relationships, enabling smooth transitions between source images and their counterfactuals. Through extensive experiments on four image datasets, we demonstrate that Mirror-CFE achieves superior performance in validity while maintaining input resemblance compared to state-of-the-art explanation methods. Finally, mirror-CFE provides interpretable visualization of the classifier's decision process by generating step-wise transitions that reveal how features evolve as classification confidence changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d74ea6a-9e10-4273-b785-431e356113b9Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald et al.ICCV 2021 · 181 citations
- Diffusion Visual Counterfactual ExplanationsMaximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias HeinNeurIPS 2022 · 124 citations
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 122 citations
Related papers
- Designing Counterfactual Generators using Deep Model InversionJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang et al.NeurIPS 2021 · 25 citations
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual ExplanationsChao Wang, chengan che, Xinyue Chen, Sophia Tsoka et al.CVPR 2026 · 1 citation
- Accurate Explanation Model for Image Classifiers using Class Association EmbeddingRuitao Xie, Jingbang Chen, Limai Jiang, Rui Xiao et al.ICDE 2024 · 12 citations
- Cycle-Consistent Counterfactuals by Latent TransformationsSaeed Khorram, Fuxin LiCVPR 2022 · 27 citations
- Grounding Counterfactual Explanation of Image Classifiers to Textual Concept SpaceSiwon Kim, Jinoh Oh, Sungjin Lee, Seunghak Yu et al.CVPR 2023
