Fooling SHAP with Stealthily Biased Sampling
Gabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand, Foutse Khomh
Abstract
SHAP explanations aim at identifying which features contribute the most to the difference in model prediction at a specific input versus a background distribution. Recent studies have shown that they can be manipulated by malicious adversaries to produce arbitrary desired explanations. However, existing attacks focus solely on altering the black-box model itself. In this paper, we propose a complementary family of attacks that leave the model intact and manipulate SHAP explanations using stealthily biased sampling of the data points used to approximate expectations w.r.t the background distribution. In the context of fairness audit, we show that our attack can reduce the importance of a sensitive feature when explaining the difference in outcomes between groups while remaining undetected. More precisely, experiments performed on real-world datasets showed that our attack could yield up to a 90% relative decrease in amplitude of the sensitive feature attribution. These results highlight the manipulability of SHAP explanations and encourage auditors to treat them with skepticism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74768eef-a92d-4460-ab82-aa05aec7a4e6Cited by top-tier papers4
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 13 citations
- Robust ML Auditing using Prior KnowledgeJade Garcia Bourrée, Augustin Godinot, Sayan Biswas, Anne-Marie Kermarrec et al.ICML 2025
- Efficient and Accurate Explanation Estimation with Distribution CompressionHubert Baniecki, Giuseppe Casalicchio, Bernd Bischl, Przemyslaw BiecekICLR 2025
Builds on1
Related papers
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
- X-Hacking: The Threat of Misguided AutoMLRahul Sharma, Sumantrak Mukherjee, Andrea Sipka, Eyke Hüllermeier et al.ICML 2025
- Unfooling Perturbation-Based Post Hoc ExplainersZachariah Carmichael, Walter J. ScheirerAAAI 2023 · 18 citations
- Washing The Unwashable : On The (Im)possibility of Fairwashing DetectionAli Shahin Shamsabadi, Mohammad Yaghini, Natalie Dullerud, Sierra Calanda Wyllie et al.NeurIPS 2022 · 23 citations
