"Why Not Other Classes?": Towards Class-Contrastive Back-Propagation Explanations
Yipei Wang, Xiaoqian Wang
摘要
Numerous methods have been developed to explain the inner mechanism of deep neural network (DNN) based classifiers. Existing explanation methods are often limited to explaining predictions of a pre-specified class, which answers the question “why is the input classified into this class?” However, such explanations with respect to a single class are inherently insufficient because they do not capture features with class-discriminative power. That is, features that are important for predicting one class may also be important for other classes. To capture features with true class-discriminative power, we should instead ask “why is the input classified into this class, but not others ?” To answer this question, we propose a weighted contrastive framework for explaining DNNs. Our framework can easily convert any existing back-propagation explanation methods to build class-contrastive explanations. We theoretically validate our weighted contrast explanation in general back-propagation explanations, and show that our framework enables class-contrastive explanations with significant improvements in both qualitative and quantitative experiments. Based on the results, we point out an important blind spot in the current explainable artificial intelligence (XAI) study, where explanations towards the predicted logits and the probabilities are obfus-cated. We suggest that these two aspects should be distinguished explicitly any time explanation methods are applied.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LICO: Explainable Models with Language-Image COnsistencyYiming Lei, Zilong Li, Yangyang Li, Junping Zhang 等NeurIPS 2023 · 被引用 12 次
- Explaining Probabilistic Models with Distributional ValuesLuca Franceschi, Michele Donini, Cédric Archambeau, Matthias W. SeegerICML 2024 · 被引用 4 次
- Benchmarking Deletion Metrics with the Principled ExplanationsYipei Wang, Xiaoqian WangICML 2024 · 被引用 4 次
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for ExplanationsYehonatan Elisha, Seffi Cohen, Oren Barkan, Noam KoenigsteinAAAI 2026 · 被引用 3 次
- Adaptive Logit Adjustment for Debiasing Multimodal Language ModelsHoin Jung, Junyi Chai, Xiaoqian WangICLR 2026
它引用的顶会 Paper4
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Fast Axiomatic Attribution for Neural NetworksRobin Hesse, Simone Schaub-Meyer, Stefan RothNeurIPS 2021 · 被引用 55 次
- A Framework to Learn with InterpretationJayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-BucNeurIPS 2021 · 被引用 35 次
- SCOUT: Self-Aware Discriminant Counterfactual ExplanationsPei Wang, Nuno VasconcelosCVPR 2020
相关 Paper
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng 等CVPR 2021
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo 等ICCV 2023 · 被引用 15 次
- Leveraging Latent Features for Local ExplanationsRonny Luss, Pin-Yu Chen, Amit Dhurandhar, Prasanna Sattigeri 等KDD 2021 · 被引用 11 次
- How to Probe: Simple Yet Effective Techniques for Improving Post-hoc ExplanationsSiddhartha Gairola, Moritz Böhle, Francesco Locatello, Bernt SchieleICLR 2025
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar 等EMNLP 2021 · 被引用 12 次
