Counterfactual Explanations Can Be Manipulated
Dylan Slack, Anna Hilgard, Himabindu Lakkaraju, Sameer Singh
摘要
Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical applications (e.g. law enforcement, financial lending), it becomes important to ensure that we clearly understand the vulnerabilties of these methods and find ways to address them. However, there is little understanding of the vulnerabilities and shortcomings of counterfactual explanations. In this work, we introduce the first framework that describes the vulnerabilities of counterfactual explanations and shows how they can be manipulated. More specifically, we show counterfactual explanations may converge to drastically different counterfactuals under a small perturbation indicating they are not robust. Leveraging this insight, we introduce a novel objective to train seemingly fair models where counterfactual explanations find much lower cost recourse under a slight perturbation. We describe how these models can unfairly provide low-cost recourse for specific subgroups in the data while appearing fair to auditors. We perform experiments on loan and violent crime prediction data sets where certain subgroups achieve up to 20x lower cost recourse under the perturbation. These results raise concerns regarding the dependability of current counterfactual explanation techniques, which we hope will inspire investigations in robust counterfactual explanations. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Framework for Evaluating Faithfulness of Local ExplanationsSanjoy Dasgupta, Nave Frost, Michal MoshkovitzICML 2022 · 被引用 87 次
- On the Adversarial Robustness of Causal Algorithmic RecourseRicardo Dominguez-Olmedo, Amir-Hossein Karimi, Bernhard SchölkopfICML 2022 · 被引用 80 次
- GLOBE-CE: A Translation Based Approach for Global Counterfactual ExplanationsDan Ley, Saumitra Mishra, Daniele MagazzeniICML 2023 · 被引用 30 次
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 被引用 28 次
- Understanding Visual Feature Reliance through the Lens of ComplexityThomas Fel, Louis Béthune, Andrew K. Lampinen, Thomas Serre 等NeurIPS 2024 · 被引用 20 次
它引用的顶会 Paper5
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approachAmir-Hossein Karimi, Bodo Julius von Kügelgen, Bernhard Schölkopf, Isabel ValeraNeurIPS 2020 · 被引用 224 次
- Predictive Multiplicity in ClassificationCharles T. Marx, Flávio P. Calmon, Berk UstunICML 2020 · 被引用 197 次
- Fairwashing explanations with off-manifold detergentChristopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller 等ICML 2020 · 被引用 104 次
- On the Fairness of Causal Algorithmic RecourseJulius von Kügelgen, Amir-Hossein Karimi, Umang Bhatt, Isabel Valera 等AAAI 2022 · 被引用 99 次
- Decisions, Counterfactual Explanations and Strategic BehaviorStratis Tsirtsis, Manuel Gomez RodriguezNeurIPS 2020 · 被引用 75 次
相关 Paper
- Learning Models for Actionable RecourseAlexis Ross, Himabindu Lakkaraju, Osbert BastaniNeurIPS 2021 · 被引用 25 次
- Probabilistically Robust Recourse: Navigating the Trade-offs between Costs and Robustness in Algorithmic RecourseMartin Pawelczyk, Teresa Datta, Johannes van den Heuvel, Gjergji Kasneci 等ICLR 2023 · 被引用 13 次
- GLANCE: Global Actions in a Nutshell for Counterfactual ExplainabilityLoukas Kavouras, Eleni Psaroudaki, Konstantinos Tsopelas, Dimitrios Rontogiannis 等AAAI 2026 · 被引用 7 次
- ElliCE: Efficient and Provably Robust Algorithmic Recourse via the Rashomon SetsBohdan Turbal, Iryna Voitsitska, Lesia SemenovaNeurIPS 2025 · 被引用 6 次
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 被引用 145 次
