"Why Not Other Classes?": Towards Class-Contrastive Back-Propagation Explanations
Yipei Wang, Xiaoqian Wang
Abstract
Numerous methods have been developed to explain the inner mechanism of deep neural network (DNN) based classifiers. Existing explanation methods are often limited to explaining predictions of a pre-specified class, which answers the question “why is the input classified into this class?” However, such explanations with respect to a single class are inherently insufficient because they do not capture features with class-discriminative power. That is, features that are important for predicting one class may also be important for other classes. To capture features with true class-discriminative power, we should instead ask “why is the input classified into this class, but not others ?” To answer this question, we propose a weighted contrastive framework for explaining DNNs. Our framework can easily convert any existing back-propagation explanation methods to build class-contrastive explanations. We theoretically validate our weighted contrast explanation in general back-propagation explanations, and show that our framework enables class-contrastive explanations with significant improvements in both qualitative and quantitative experiments. Based on the results, we point out an important blind spot in the current explainable artificial intelligence (XAI) study, where explanations towards the predicted logits and the probabilities are obfus-cated. We suggest that these two aspects should be distinguished explicitly any time explanation methods are applied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- LICO: Explainable Models with Language-Image COnsistencyYiming Lei, Zilong Li, Yangyang Li, Junping Zhang et al.NeurIPS 2023 · 12 citations
- Explaining Probabilistic Models with Distributional ValuesLuca Franceschi, Michele Donini, Cédric Archambeau, Matthias W. SeegerICML 2024 · 4 citations
- Benchmarking Deletion Metrics with the Principled ExplanationsYipei Wang, Xiaoqian WangICML 2024 · 4 citations
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for ExplanationsYehonatan Elisha, Seffi Cohen, Oren Barkan, Noam KoenigsteinAAAI 2026 · 3 citations
- Adaptive Logit Adjustment for Debiasing Multimodal Language ModelsHoin Jung, Junyi Chai, Xiaoqian WangICLR 2026
Builds on4
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Fast Axiomatic Attribution for Neural NetworksRobin Hesse, Simone Schaub-Meyer, Stefan RothNeurIPS 2021 · 55 citations
- A Framework to Learn with InterpretationJayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-BucNeurIPS 2021 · 35 citations
- SCOUT: Self-Aware Discriminant Counterfactual ExplanationsPei Wang, Nuno VasconcelosCVPR 2020
Related papers
- A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsYunhao Ge, Yao Xiao, Zhi Xu, Meng Zheng et al.CVPR 2021
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo et al.ICCV 2023 · 15 citations
- Leveraging Latent Features for Local ExplanationsRonny Luss, Pin-Yu Chen, Amit Dhurandhar, Prasanna Sattigeri et al.KDD 2021 · 11 citations
- How to Probe: Simple Yet Effective Techniques for Improving Post-hoc ExplanationsSiddhartha Gairola, Moritz Böhle, Francesco Locatello, Bernt SchieleICLR 2025
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar et al.EMNLP 2021 · 12 citations
