Making the Classification Explanation Faithful to the Confidence Score
Jian-Xun Mi, Lu Pan, Weisheng Li
摘要
Deep Neural Networks have revolutionized numerous industries, yet their decision-making processes remain largely opaque. Most existing explanation methods visualize the importance of image regions that influence a classifier's decisions, but they predominantly focus on identifying regions with positive contributions, often overlooking those with negative impacts. In this paper, we introduce a novel black-box explanation method, the Metropolis-Hastings Explainer (MHE), designed to provide confidence-faithful explanations. MHE enhances the fidelity of explanations by ensuring that the explained regions closely align with the original confidence score, sampling instances that best match the classifier's confidence. Furthermore, MHE improves sampling efficiency by utilizing existing valid samples to explore more potential valid ones, reducing computational overhead. To enhance the clarity of explanations, MHE prioritizes valid samples with smaller areas when other factors are equal, thereby reducing the explanation area. Building upon the MHE framework, we propose two extensions: MHE-e, which focuses exclusively on regions with positive contributions, and MHE-pro, which refines explanation quality by integrating multi-scale information. MHE-pro progressively regions, optimizing both sampling efficiency and explanation quality. Experimental results demonstrate that MHE delivers superior and stable explanation quality across various models, including ResNet50, VGG16, ViT, DINO, and CLIP, on datasets such as ImageNet, CUB-200-2011, and VOC2012, providing explanations that closely approximate the original classification confidence. The source code and demo are available at https://github.com/helloAI007/MHE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 被引用 147 次
相关 Paper
- What You See is What You Classify: Black Box AttributionsSteven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz 等NeurIPS 2022 · 被引用 15 次
- MaskDiME: Adaptive Masked Diffusion for Precise and Efficient Visual Counterfactual ExplanationsChanglu Guo, Anders Nymark Christensen, Anders Bjorholm Dahl, Morten Rieger HannemoseCVPR 2026 · 被引用 2 次
- Salvage: Shapley-distribution Approximation Learning Via Attribution Guided Exploration for Explainable Image ClassificationMehdi Naouar, Hanne Raum, Jens Rahnfeld, Yannick Vogt 等ICLR 2025
- Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsTownim Faisal Chowdhury, Vu Minh Hieu Phan, Kewen Liao, Nanyu Dong 等ICCV 2025
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 被引用 46 次
