Fairwashing explanations with off-manifold detergent
Christopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, Pan Kessel
Abstract
Explanation methods promise to make black-box classifiers more transparent. As a result, it is hoped that they can act as proof for a sensible, fair and trustworthy decision-making process of the algorithm and thereby increase its acceptance by the end-users. In this paper, we show both theoretically and experimentally that these hopes are presently unfounded. Specifically, we show that, for any classifier , one can always construct another classifier which has the same behavior on the data (same train, validation, and test error) but has arbitrarily manipulated explanation maps. We derive this statement theoretically using differential geometry and demonstrate it experimentally for various explanation methods, architectures, and datasets. Motivated by our theoretical insights, we then propose a modification of existing explanation methods which makes them significantly more robust.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd0deb97-3447-4334-b58a-39f1420fbe10Cited by top-tier papers23
- Explainable Deep One-Class ClassificationPhilipp Liznerski, Lukas Ruff, Robert A. Vandermeulen, Billy Joe Franks et al.ICLR 2021 · 240 citations
- Counterfactual Explanations Can Be ManipulatedDylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer SinghNeurIPS 2021 · 182 citations
- Manifold Preserving Guided DiffusionYutong He, Naoki Murata, Chieh-Hsin Lai, Yuhta Takida et al.ICLR 2024 · 148 citations
- Shapley explainability on the data manifoldChristopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton et al.ICLR 2021 · 125 citations
- Towards Multi-Grained Explainability for Graph Neural NetworksXiang Wang, Ying-Xin Wu, An Zhang, Xiangnan He et al.NeurIPS 2021 · 105 citations
Related papers
- Corrupting Neuron Explanations of Deep Visual FeaturesDivyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei WengICCV 2023 · 3 citations
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand et al.ICLR 2023 · 3 citations
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Fooling Explanations in Text ClassifiersAdam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal FrossardICLR 2022 · 23 citations
