Characterizing the risk of fairwashing
Ulrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi Hara
摘要
Fairwashing refers to the risk that an unfair black-box model can be explained by a fairer model through post-hoc explanation manipulation. In this paper, we investigate the capability of fairwashing attacks by analyzing their fidelity-unfairness trade-offs. In particular, we show that fairwashed explanation models can generalize beyond the suing group (i.e., data points that are being explained), meaning that a fairwashed explainer can be used to rationalize subsequent unfair decisions of a black-box model. We also demonstrate that fairwashing attacks can transfer across black-box models, meaning that other black-box models can perform fairwashing without explicitly using their predictions. This generalization and transferability of fairwashing attacks imply that their detection will be difficult in practice. Finally, we propose an approach to quantify the risk of fairwashing, which is based on the computation of the range of the unfairness of high-fidelity explainers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 被引用 41 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Washing The Unwashable : On The (Im)possibility of Fairwashing DetectionAli Shahin Shamsabadi, Mohammad Yaghini, Natalie Dullerud, Sierra Calanda Wyllie 等NeurIPS 2022 · 被引用 23 次
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg 等AAAI 2024 · 被引用 13 次
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr 等NeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper4
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 被引用 98 次
- Ensuring Fairness Beyond the Training DataDebmalya Mandal, Samuel Deng, Suman Jana, Jeannette M. Wing 等NeurIPS 2020 · 被引用 68 次
相关 Paper
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand 等ICLR 2023 · 被引用 3 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
- A Theory of Transfer-Based Black-Box Attacks: Explanation and ImplicationsYanbo Chen, Weiwei LiuNeurIPS 2023 · 被引用 22 次
- Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box DomainsQilong Zhang, Xiaodan Li, Yuefeng Chen, Jingkuan Song 等ICLR 2022 · 被引用 85 次
