Explaining Black-Box Algorithms Using Probabilistic Contrastive Counterfactuals
Sainyam Galhotra, Romila Pradhan, Babak Salimi
摘要
There has been a recent resurgence of interest in explainable artificial intelligence (XAI) that aims to reduce the opaqueness of AI-based decision-making systems, allowing humans to scrunitize and trust them. Prior work in this context has focused on the attribution of responsibility for an algorithm's decisions to its inputs wherein responsibility is typically approached as a purely associational concept. In this paper, we propose a principled causality-based approach for explaining black-box decision-making systems that addresses limitations of existing methods in XAI. At the core of our framework lies probabilistic contrastive counterfactuals, a concept that can be traced back to philosophical, cognitive, and social foundations of theories on how humans generate and select explanations. We show how such counterfactuals can quantify the direct and indirect influences of a variable on decisions made by an algorithm, and provide actionable recourse for individuals negatively affected by the algorithm's decision. Unlike prior work, our system, LEWIS: (1) can compute provably effective explanations and recourse at local, global and contextual levels; (2) is designed to work with users with varying levels of background knowledge of the underlying causal model; and (3) makes no assumptions about the internals of an algorithmic system except for the availability of its input-output data. We empirically evaluate LEWIS on three real-world datasets and show that it generates human-understandable explanations that improve upon state-of-the-art approaches in XAI, including the popular LIME and SHAP. Experiments on synthetic data further demonstrate the correctness of LEWIS's explanations and the scalability of its recourse algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 被引用 53 次
- Towards Trustworthy Explanation: On Causal RationalizationWenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai 等ICML 2023 · 被引用 25 次
- XInsight: eXplainable Data Analysis Through The Lens of CausalityPingchuan Ma, Rui Ding, Shuai Wang, Shi Han 等SIGMOD 2023 · 被引用 20 次
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 被引用 19 次
- Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningXiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li 等NeurIPS 2025 · 被引用 19 次
它引用的顶会 Paper6
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 被引用 774 次
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 被引用 458 次
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 被引用 246 次
- Algorithmic recourse under imperfect causal knowledge: a probabilistic approachAmir-Hossein Karimi, Bodo Julius von Kügelgen, Bernhard Schölkopf, Isabel ValeraNeurIPS 2020 · 被引用 224 次
- Strategic Classification is Causal Modeling in DisguiseJohn Miller, Smitha Milli, Moritz HardtICML 2020 · 被引用 127 次
相关 Paper
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable RecoursesKaivalya Rawal, Himabindu LakkarajuNeurIPS 2020 · 被引用 113 次
- Feature Responsiveness Scores: Model-Agnostic Explanations for RecourseSeung Hyun Cheon, Anneke Wernerfelt, Sorelle A. Friedler, Berk UstunICLR 2025
- Let the CAT out of the bag: Contrastive Attributed explanations for TextSaneem A. Chemmengath, Amar Prakash Azad, Ronny Luss, Amit DhurandharEMNLP 2022 · 被引用 6 次
- Learning Feasible Causal Algorithmic Recourse: A Prior Structural Knowledge Free ApproachHaotian Wang, Hao Zou, Xueguang Zhou, Shangwen Wang 等WWW 2025 · 被引用 1 次
- Generative Perturbation Analysis for Probabilistic Black-Box Anomaly AttributionTsuyoshi Idé, Naoki AbeKDD 2023 · 被引用 4 次
