Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural Networks
Xue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo, Zhifei Zhang, Lu Jin, Tao Wei, Kui Ren
摘要
Explaining deep models in a human-understandable way has been explored by many works that mostly explain why an input causes a corresponding prediction (i.e., Why P?). However, seldom they could handle those more complex causal questions like "Why P rather than Q?" and "Why one is P, while another is Q?", which would better help humans understand the behavior of deep models. Considering the insufficient study on such complex causal questions, we make the first attempt to explain different causal questions by contrastive explanations in a unified framework, i.e., Counterfactual Contrastive Explanation (CCE), which visually and intuitively explains the aforementioned questions via a novel positive-negative saliency-based explanation scheme. More specifically, we propose a content-aware counterfactual perturbing algorithm to stimulate contrastive examples, from which a pair of positive and negative saliency maps could be derived to contrastively explain why P (positive class) rather than Q (negative class). Beyond existing works, our counterfactual perturbation meets the principles of validity, sparsity, and data distribution closeness at the same time. In addition, by slightly adjusting the objective of perturbation, our framework can adapt to different causal questions. Extensive experimental evaluation demonstrates the effectiveness and superior performance of the proposed CCE on different benchmark metrics for interpretability, including Sanity Check, Class Deviation Score and Insertion-Deletion tests. A user study is conducted and the results show that user confidence is increasing significantly when presented with CCE compared to standard saliency map baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful ReasoningTianrun Xu, Haoda Jing, Ye Li, Yuquan Wei 等ICML 2026 · 被引用 8 次
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for ExplanationsYehonatan Elisha, Seffi Cohen, Oren Barkan, Noam KoenigsteinAAAI 2026 · 被引用 3 次
- What is Missing? Explaining Neurons Activated by Absent ConceptsRobin Hesse, Simone Schaub-Meyer, Janina Hesse, Bernt Schiele 等ICML 2026 · 被引用 1 次
- Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency PartitionLintong Zhang, Kang Yin, Seong-Whan LeeCVPR 2025
- Image-based Outlier Synthesis With Training DataSudarshan RegmiCVPR 2026
它引用的顶会 Paper10
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Feature Importance-aware Transferable Adversarial AttacksZhibo Wang, Hengchang Guo, Zhifei Zhang, Wenxin Liu 等ICCV 2021 · 被引用 306 次
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 被引用 251 次
- Visualizing Deep Networks by Optimizing with Integrated GradientsZhongang Qi, Saeed Khorram, Fuxin LiAAAI 2020 · 被引用 149 次
- CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-LinesArjun R. Akula, Shuai Wang, Song-Chun ZhuAAAI 2020 · 被引用 102 次
相关 Paper
- KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language InferenceQianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li 等ACL 2021
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li 等AAAI 2024 · 被引用 12 次
- Counterfactual Concept Bottleneck ModelsGabriele Dominici, Pietro Barbiero, Francesco Giannini, Martin Gjoreski 等ICLR 2025
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 被引用 122 次
- "Why Not Other Classes?": Towards Class-Contrastive Back-Propagation ExplanationsYipei Wang, Xiaoqian WangNeurIPS 2022 · 被引用 17 次
