A Consistent and Efficient Evaluation Strategy for Attribution Methods
Yao Rong, Tobias Leemann, Vadim Borisov, Gjergji Kasneci, Enkelejda Kasneci
Abstract
With a variety of local feature attribution methods being proposed in recent years, follow-up work suggested several evaluation strategies. To assess the attribution quality across different attribution techniques, the most popular among these evaluation strategies in the image domain use pixel perturbations. However, recent advances discovered that different evaluation strategies produce conflicting rankings of attribution methods and can be prohibitively expensive to compute. In this work, we present an informationtheoretic analysis of evaluation strategies based on pixel perturbations. Our findings reveal that the results are strongly affected by information leakage through the shape of the removed pixels as opposed to their actual values. Using our theoretical insights, we propose a novel evaluation framework termed Remove and Debias (ROAD) which offers two contributions: First, it mitigates the impact of the confounders, which entails higher consistency among evaluation strategies. Second, ROAD does not require the computationally expensive retraining step and saves up to 99 % in computational costs compared to the state-of-the-art. We release our source code at https://github.com/ tleemann/road_evaluation .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck AttributionYing Wang, Tim G. J. Rudner, Andrew Gordon WilsonNeurIPS 2023 · 51 citations
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 18 citations
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- Probabilistic Stability Guarantees for Feature AttributionsHelen Jin, Anton Xue, Weiqiu You, Surbhi Goel et al.NeurIPS 2025 · 12 citations
- On the Faithfulness of Vision Transformer ExplanationsJunyi Wu, Weitai Kang, Hao Tang, Yuan Hong et al.CVPR 2024 · 8 citations
Builds on6
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
Related papers
- Pixel-level Certified Explanations via Randomized SmoothingAlaa Anani, Tobias Lorenz, Mario Fritz, Bernt SchieleICML 2025
- F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AIXu Zheng, Farhad Shirani, Zhuomin Chen, Chaohao Lin et al.ICLR 2025
- Faithfulness Under the Distribution: A New Look at Attribution EvaluationZhiyu Zhu, Zhibo Jin, Jiayu Zhang, Bartlomiej Sobieski et al.ICLR 2026
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
