Sanity Simulations for Saliency Methods
Joon Sik Kim, Gregory Plumb, Ameet Talwalkar
Abstract
Saliency methods are a popular class of feature attribution explanation methods that aim to capture a model's predictive reasoning by identifying "important" pixels in an input image. However, the development and adoption of these methods are hindered by the lack of access to ground-truth model reasoning, which prevents accurate evaluation. In this work, we design a synthetic benchmarking framework, SMERF, that allows us to perform ground-truth-based evaluation while controlling the complexity of the model's reasoning. Experimentally, SMERF reveals significant limitations in existing saliency methods and, as a result, represents a useful tool for the development of new saliency methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3bf051e-3559-4774-ae84-b55d3d8d7fb5Cited by top-tier papers2
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- FaCT: Faithful Concept Traces for Explaining Neural Network DecisionsAmin Parchami-Araghi, Sukrut Rao, Jonas Fischer, Bernt SchieleNeurIPS 2025 · 1 citation
Builds on5
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 167 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Do Users Benefit From Interpretable Vision? A User Study, Baseline, And DatasetLeon Sixt, Martin Schuessler, Oana-Iuliana Popescu, Philipp Weiß et al.ICLR 2022 · 21 citations
Related papers
- Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model BehaviorAngie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik StrobeltCHI 2022 · 51 citations
- Rethinking Saliency Maps: A Cognitive Human Aligned Taxonomy and Evaluation Framework for ExplanationsYehonatan Elisha, Seffi Cohen, Oren Barkan, Noam KoenigsteinAAAI 2026 · 3 citations
- XRAI: Better Attributions Through RegionsAndrei Kapishnikov, Tolga Bolukbasi, Fernanda B. Viégas, Michael TerryICCV 2019 · 251 citations
- New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and SoundArushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu et al.NeurIPS 2022 · 12 citations
- Red Teaming Deep Neural Networks with Feature Synthesis ToolsStephen Casper, Tong Bu, Yuxiao Li, Jiawei Li et al.NeurIPS 2023 · 23 citations
