Washing The Unwashable : On The (Im)possibility of Fairwashing Detection
Ali Shahin Shamsabadi, Mohammad Yaghini, Natalie Dullerud, Sierra Calanda Wyllie, Ulrich Aïvodji, Aisha Alaagib, Sébastien Gambs, Nicolas Papernot
Abstract
The use of black-box models (e.g., deep neural networks) in high-stakes decisionmaking systems, whose internal logic is complex, raises the need for providing explanations about their decisions. Model explanation techniques mitigate this problem by generating an interpretable and high-fidelity surrogate model (e.g., a logistic regressor or decision tree) to explain the logic of black-box models. In this work, we investigate the issue of fairwashing, in which model explanation techniques are manipulated to rationalize decisions taken by an unfair black-box model using deceptive surrogate models. More precisely, we theoretically characterize and analyze fairwashing, proving that this phenomenon is difficult to avoid due to an irreducible factor-the unfairness of the black-box model. Based on the theory developed, we propose a novel technique, called FRAUD-Detect (FaiRness AUDit Detection), to detect fairwashed models by measuring a divergence over subpopulation-wise fidelity measures of the interpretable model. We empirically demonstrate that this divergence is significantly larger in purposefully fairwashed interpretable models than in honest ones. Furthermore, we show that our detector is robust to an informed adversary trying to bypass our detector. The code implementing FRAUD-Detect is available at https://github.com/cleverhans-lab/FRAUD-Detect . ⇤ Contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fdc2bb2-828b-4a7c-a677-469cd8038beeCited by top-tier papers7
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 41 citations
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel et al.ICML 2024 · 38 citations
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr et al.NeurIPS 2024 · 10 citations
- Secure and Confidential Certificates of Online FairnessOlive Franzese, Ali Shahin Shamsabadi, Carter Luck, Hamed HaddadiNeurIPS 2025 · 10 citations
- Robust ML Auditing using Prior KnowledgeJade Garcia Bourrée, Augustin Godinot, Sayan Biswas, Anne-Marie Kermarrec et al.ICML 2025
Builds on4
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 197 citations
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller et al.ICML 2020 · 104 citations
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 98 citations
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
Related papers
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand et al.ICLR 2023 · 3 citations
- Unfooling Perturbation-Based Post Hoc ExplainersZachariah Carmichael, Walter J. ScheirerAAAI 2023 · 18 citations
- Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated ContentHaoyue Bai, Yiyou Sun, Wei Cheng, Haifeng ChenCVPR 2025
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun et al.ICML 2026
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg et al.AAAI 2024 · 13 citations
