Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
Nils Philipp Walter, Jilles Vreeken, Jonas Fischer
Abstract
Attribution methods reveal which input features a neural network uses for a prediction, adding transparency to their decisions. A common problem is that these attributions seem unspecific, highlighting both important and irrelevant features. We revisit the common attribution pipeline and observe that using logits as attribution target is a main cause of this phenomenon. We show that the solution is in plain sight: considering distributions of attributions over multiple classes using existing attribution methods yields specific and fine-grained attributions. On common benchmarks, including the grid-pointing game and randomization-based sanity checks, this improves the ability of 18 attribution methods across 7 architectures up to , agnostic to model architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f847f21-7482-487b-b3b2-7895fef4fd01Cited by top-tier papers1
Ask how each one uses itBuilds on9
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 62 citations
- A Framework to Learn with InterpretationJayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-BucNeurIPS 2021 · 35 citations
- Towards Better Understanding Attribution MethodsSukrut Rao, Moritz Böhle, Bernt SchieleCVPR 2022 · 32 citations
Related papers
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 167 citations
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov et al.ICLR 2026 · 8 citations
- DAVE: Distribution-aware Attribution via ViT Gradient DecompositionAdam Wróbel, Siddhartha Gairola, Jacek Tabor, Bernt Schiele et al.ICML 2026 · 2 citations
- SAM: The Sensitivity of Attribution Methods to HyperparametersNaman Bansal, Chirag Agarwal, Anh NguyenCVPR 2020
