Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification Models
Townim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong, Minh-Son To, Anton van den Hengel, Johan W. Verjans, Zhibin Liao
摘要
Counterfactual explanations (CFE) for deep image classifiers aim to reveal how minimal input changes lead to different model decisions, providing critical insights for model interpretation and improvement. However, existing CFE methods often rely on additional image encoders and generative models to create plausible images, neglecting the classifier's own feature space and decision boundaries. As such, they do not explain the intrinsic feature space and decision boundaries learned by the classifier. To address this limitation, we propose Mirror-CFE, a novel method that generates faithful counterfactual explanations by operating directly in the classifier's feature space, treating decision boundaries as mirrors that ``reflect'' feature representations in the mirror. Mirror-CFE learns a mapping function from feature space to image space while preserving distance relationships, enabling smooth transitions between source images and their counterfactuals. Through extensive experiments on four image datasets, we demonstrate that Mirror-CFE achieves superior performance in validity while maintaining input resemblance compared to state-of-the-art explanation methods. Finally, mirror-CFE provides interpretable visualization of the classifier's decision process by generating step-wise transitions that reveal how features evolve as classification confidence changes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald 等ICCV 2021 · 被引用 181 次
- Diffusion Visual Counterfactual ExplanationsMaximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias HeinNeurIPS 2022 · 被引用 124 次
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 被引用 122 次
相关 Paper
- Designing Counterfactual Generators using Deep Model InversionJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang 等NeurIPS 2021 · 被引用 25 次
- Back to the Feature: Explaining Video Classifiers with Video Counterfactual ExplanationsChao Wang, chengan che, Xinyue Chen, Sophia Tsoka 等CVPR 2026 · 被引用 1 次
- Accurate Explanation Model for Image Classifiers using Class Association EmbeddingRuitao Xie, Jingbang Chen, Limai Jiang, Rui Xiao 等ICDE 2024 · 被引用 12 次
- Cycle-Consistent Counterfactuals by Latent TransformationsSaeed Khorram, Fuxin LiCVPR 2022 · 被引用 27 次
- Grounding Counterfactual Explanation of Image Classifiers to Textual Concept SpaceSiwon Kim, Jinoh Oh, Sungjin Lee, Seunghak Yu 等CVPR 2023
