Faithful Model Explanations through Energy-Constrained Conformal Counterfactuals
Patrick Altmeyer, Mojtaba Farmanbar, Arie van Deursen, Cynthia C. S. Liem
摘要
Counterfactual explanations offer an intuitive and straightforward way to explain black-box models and offer algorithmic recourse to individuals. To address the need for plausible explanations, existing work has primarily relied on surrogate models to learn how the input data is distributed. This effectively reallocates the task of learning realistic explanations for the data from the model itself to the surrogate. Consequently, the generated explanations may seem plausible to humans but need not necessarily describe the behaviour of the black-box model faithfully. We formalise this notion of faithfulness through the introduction of a tailored evaluation metric and propose a novel algorithmic framework for generating Energy-Constrained Conformal Counterfactuals that are only as plausible as the model permits. Through extensive empirical studies, we demonstrate that ECCCo reconciles the need for faithfulness and plausibility. In particular, we show that for models with gradient access, it is possible to achieve state-of-the-art performance without the need for surrogate models. To do so, our framework relies solely on properties defining the black-box model itself by leveraging recent advances in energy-based modelling and conformal prediction. To our knowledge, this is the first venture in this direction for generating faithful counterfactual explanations. Thus, we anticipate that ECCCo can serve as a baseline for future research. We believe that our work opens avenues for researchers and practitioners seeking tools to better distinguish trustworthy from unreliable models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 被引用 145 次
- Probabilistically Robust Recourse: Navigating the Trade-offs between Costs and Robustness in Algorithmic RecourseMartin Pawelczyk, Teresa Datta, Johannes van den Heuvel, Gjergji Kasneci 等ICLR 2023 · 被引用 13 次
相关 Paper
- Learning Global Transparent Models consistent with Local Contrastive ExplanationsTejaswini Pedapati, Avinash Balakrishnan, Karthikeyan Shanmugam, Amit DhurandharNeurIPS 2020 · 被引用 35 次
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 被引用 122 次
- Explaining Black-Box Algorithms Using Probabilistic Contrastive CounterfactualsSainyam Galhotra, Romila Pradhan, Babak SalimiSIGMOD 2021 · 被引用 85 次
- Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIWon Jun Kim, Hyungjin Chung, Jaemin Kim, Sangmin Lee 等CVPR 2025
- Learning Feasible Causal Algorithmic Recourse: A Prior Structural Knowledge Free ApproachHaotian Wang, Hao Zou, Xueguang Zhou, Shangwen Wang 等WWW 2025 · 被引用 1 次
