Characterizing the risk of fairwashing
Ulrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi Hara
Abstract
Fairwashing refers to the risk that an unfair black-box model can be explained by a fairer model through post-hoc explanation manipulation. In this paper, we investigate the capability of fairwashing attacks by analyzing their fidelity-unfairness trade-offs. In particular, we show that fairwashed explanation models can generalize beyond the suing group (i.e., data points that are being explained), meaning that a fairwashed explainer can be used to rationalize subsequent unfair decisions of a black-box model. We also demonstrate that fairwashing attacks can transfer across black-box models, meaning that other black-box models can perform fairwashing without explicitly using their predictions. This generalization and transferability of fairwashing attacks imply that their detection will be difficult in practice. Finally, we propose an approach to quantify the risk of fairwashing, which is based on the computation of the range of the unfairness of high-fidelity explainers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fc20ed0-1010-4482-95d1-488cd3d3812bCited by top-tier papers7
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 41 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Washing The Unwashable : On The (Im)possibility of Fairwashing DetectionAli Shahin Shamsabadi, Mohammad Yaghini, Natalie Dullerud, Sierra Calanda Wyllie et al.NeurIPS 2022 · 23 citations
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg et al.AAAI 2024 · 13 citations
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr et al.NeurIPS 2024 · 10 citations
Builds on4
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 197 citations
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 98 citations
- Ensuring Fairness Beyond the Training DataDebmalya Mandal, Samuel Deng, Suman Jana, Jeannette M. Wing et al.NeurIPS 2020 · 68 citations
Related papers
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand et al.ICLR 2023 · 3 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
- A Theory of Transfer-Based Black-Box Attacks: Explanation and ImplicationsYanbo Chen, Weiwei LiuNeurIPS 2023 · 22 citations
- Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box DomainsQilong Zhang, Xiaodan Li, Yuefeng Chen, Jingkuan Song et al.ICLR 2022 · 85 citations
