Washing The Unwashable : On The (Im)possibility of Fairwashing Detection
Ali Shahin Shamsabadi, Mohammad Yaghini, Natalie Dullerud, Sierra Calanda Wyllie, Ulrich Aïvodji, Aisha Alaagib, Sébastien Gambs, Nicolas Papernot
摘要
The use of black-box models (e.g., deep neural networks) in high-stakes decisionmaking systems, whose internal logic is complex, raises the need for providing explanations about their decisions. Model explanation techniques mitigate this problem by generating an interpretable and high-fidelity surrogate model (e.g., a logistic regressor or decision tree) to explain the logic of black-box models. In this work, we investigate the issue of fairwashing, in which model explanation techniques are manipulated to rationalize decisions taken by an unfair black-box model using deceptive surrogate models. More precisely, we theoretically characterize and analyze fairwashing, proving that this phenomenon is difficult to avoid due to an irreducible factor-the unfairness of the black-box model. Based on the theory developed, we propose a novel technique, called FRAUD-Detect (FaiRness AUDit Detection), to detect fairwashed models by measuring a divergence over subpopulation-wise fidelity measures of the interpretable model. We empirically demonstrate that this divergence is significantly larger in purposefully fairwashed interpretable models than in honest ones. Furthermore, we show that our detector is robust to an informed adversary trying to bypass our detector. The code implementing FRAUD-Detect is available at https://github.com/cleverhans-lab/FRAUD-Detect . ⇤ Contributed equally.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 被引用 41 次
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel 等ICML 2024 · 被引用 38 次
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr 等NeurIPS 2024 · 被引用 10 次
- Secure and Confidential Certificates of Online FairnessOlive Franzese, Ali Shahin Shamsabadi, Carter Luck, Hamed HaddadiNeurIPS 2025 · 被引用 10 次
- Robust ML Auditing using Prior KnowledgeJade Garcia Bourrée, Augustin Godinot, Sayan Biswas, Anne-Marie Kermarrec 等ICML 2025
它引用的顶会 Paper4
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
- Characterizing Fairness Over the Set of Good Models Under Selective LabelsAmanda Coston, Ashesh Rambachan, Alexandra ChouldechovaICML 2021 · 被引用 98 次
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 被引用 35 次
相关 Paper
- Fooling SHAP with Stealthily Biased SamplingGabriel Laberge, Ulrich Aïvodji, Satoshi Hara, Mario Marchand 等ICLR 2023 · 被引用 3 次
- Unfooling Perturbation-Based Post Hoc ExplainersZachariah Carmichael, Walter J. ScheirerAAAI 2023 · 被引用 18 次
- Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated ContentHaoyue Bai, Yiyou Sun, Wei Cheng, Haifeng ChenCVPR 2025
- RECAST: Model Reconstruction via Counterfactual-Aware Wasserstein Geometry under Limited DataXuan Zhao, Lena Krieger, Zhuo Cao, Arya Bangun 等ICML 2026
- Adversarial Attacks on the Interpretation of Neuron Activation MaximizationGéraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg 等AAAI 2024 · 被引用 13 次
