What You See is What You Classify: Black Box Attributions
Steven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz, Michele Volpi
Abstract
An important step towards explaining deep image classifiers lies in the identification of image regions that contribute to individual class scores in the model's output. However, doing this accurately is a difficult task due to the black-box nature of such networks. Most existing approaches find such attributions either using activations and gradients or by repeatedly perturbing the input. We instead address this challenge by training a second deep network, the Explainer, to predict attributions for a pre-trained black-box classifier, the Explanandum. These attributions are provided in the form of masks that only show the classifier-relevant parts of an image, masking out the rest. Our approach produces sharper and more boundaryprecise masks when compared to the saliency maps generated by other methods. Moreover, unlike most existing approaches, ours is capable of directly generating very distinct class-specific masks in a single forward pass. This makes the proposed method very efficient during inference. We show that our attributions are superior to established methods both visually and quantitatively with respect to the PASCAL VOC-2007 and Microsoft COCO-2014 datasets. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbbb3531-567b-437d-8961-bafc1c8cf4f3Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Black-Box Explanation of Object Detectors via Saliency MapsVitali Petsiuk, Rajiv Jain, Varun Manjunatha, Vlad I. Morariu et al.CVPR 2021
Related papers
- Explanations for Occluded ImagesHana Chockler, Daniel Kroening, Youcheng SunICCV 2021 · 23 citations
- Saliency is a Possible Red Herring When Diagnosing Poor GeneralizationJoseph D. Viviano, Becks Simpson, Francis Dutil, Yoshua Bengio et al.ICLR 2021 · 46 citations
- Explainable Models with Consistent InterpretationsVipin Pillai, Hamed PirsiavashAAAI 2021 · 46 citations
- Salvage: Shapley-distribution Approximation Learning Via Attribution Guided Exploration for Explainable Image ClassificationMehdi Naouar, Hanne Raum, Jens Rahnfeld, Yannick Vogt et al.ICLR 2025
- Accurate Explanation Model for Image Classifiers using Class Association EmbeddingRuitao Xie, Jingbang Chen, Limai Jiang, Rui Xiao et al.ICDE 2024 · 12 citations
