Explainable Models with Consistent Interpretations
Vipin Pillai, Hamed Pirsiavash
Abstract
Given the widespread deployment of black box deep neural networks in computer vision applications, the interpretability aspect of these black box systems has recently gained traction. Various methods have been proposed to explain the results of such deep neural networks. However, some recent works have shown that such explanation methods are biased and do not produce consistent interpretations. Hence, rather than introducing a novel explanation method, we learn models that are encouraged to be interpretable given an explanation method. We use Grad-CAM as the explanation algorithm and encourage the network to learn consistent interpretations along with maximizing the log-likelihood of the correct class. We show that our method outperforms the baseline on the pointing game evaluation on ImageNet and MS-COCO datasets respectively. We also introduce new evaluation metrics that penalize the saliency map if it lies outside the ground truth bounding box or segmentation mask, and show that our method outperforms the baseline on these metrics as well. Moreover, our model trained with interpretation consistency generalizes to other explanation algorithms on all the evaluation metrics. The code and models are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03643b51-e29f-49ef-9e74-d27730f498f3Cited by top-tier papers7
- Learning Support and Trivial Prototypes for Interpretable Image ClassificationChong Wang, Yuyuan Liu, Yuanhong Chen, Fengbei Liu et al.ICCV 2023 · 50 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- Saliency-Aware Neural Architecture SearchRamtin Hosseini, Pengtao XieNeurIPS 2022 · 16 citations
- Consistent Explanations by Contrastive LearningVipin Pillai, Soroush Abbasi Koohpayegani, Ashley Ouligian, Dennis Fong et al.CVPR 2022 · 15 citations
- B-cosification: Transforming Deep Neural Networks to be Inherently InterpretableShreyash Arya, Sukrut Rao, Moritz Böhle, Bernt SchieleNeurIPS 2024 · 14 citations
Builds on4
- Fooling Network Interpretation in Image ClassificationAkshayvarun Subramanya, Vipin Pillai, Hamed PirsiavashICCV 2019 · 68 citations
- Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic SegmentationYude Wang, Jie Zhang, Meina Kan, Shiguang Shan et al.CVPR 2020
- Don't Judge an Object by Its Context: Learning to Overcome Contextual BiasKrishna Kumar Singh, Dhruv Mahajan, Kristen Grauman, Yong Jae Lee et al.CVPR 2020
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- What You See is What You Classify: Black Box AttributionsSteven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz et al.NeurIPS 2022 · 15 citations
- Learning Global Transparent Models consistent with Local Contrastive ExplanationsTejaswini Pedapati, Avinash Balakrishnan, Karthikeyan Shanmugam, Amit DhurandharNeurIPS 2020 · 35 citations
- LICO: Explainable Models with Language-Image COnsistencyYiming Lei, Zilong Li, Yangyang Li, Junping Zhang et al.NeurIPS 2023 · 12 citations
- A Novel Visual Interpretability for Deep Neural Networks by Optimizing Activation Maps with PerturbationQing-Long Zhang, Lu Rao, Yubin YangAAAI 2021 · 26 citations
- Eye into AI: Evaluating the Interpretability of Explainable AI Techniques through a Game with a PurposeKatelyn Morrison, Mayank Jain, Jessica Hammer, Adam PererCSCW 2023 · 11 citations
