Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
Dylan Slack, Anna Hilgard, Sameer Singh, Himabindu Lakkaraju
摘要
As black box explanations are increasingly being employed to establish model credibility in high stakes settings, it is important to ensure that these explanations are accurate and reliable. However, prior work demonstrates that explanations generated by state-of-the-art techniques are inconsistent, unstable, and provide very little insight into their correctness and reliability. In addition, these methods are also computationally inefficient, and require significant hyper-parameter tuning. In this paper, we address the aforementioned challenges by developing a novel Bayesian framework for generating local explanations along with their associated uncertainty. We instantiate this framework to obtain Bayesian versions of LIME and KernelSHAP which output credible intervals for the feature importances, capturing the associated uncertainty. The resulting explanations not only enable us to make concrete inferences about their quality (e.g., there is a 95% chance that the feature importance lies within the given range), but are also highly consistent and stable. We carry out a detailed theoretical analysis that leverages the aforementioned uncertainty to estimate how many perturbations to sample, and how to sample for faster convergence. This work makes the first attempt at addressing several critical issues with popular explanation methods in one shot, thereby generating consistent, stable, and reliable explanations with guarantees in a computationally efficient manner. Experimental evaluation with multiple real world datasets and user studies demonstrate that the efficacy of the proposed framework. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health CareVenkatesh Sivaraman, Leigh A. Bukowski, Joel Levin, Jeremy M. Kahn 等CHI 2023 · 被引用 126 次
- A Holistic Approach to Unifying Automatic Concept Extraction and Concept Importance EstimationThomas Fel, Victor Boutin, Louis Béthune, Rémi Cadène 等NeurIPS 2023 · 被引用 125 次
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun 等NeurIPS 2023 · 被引用 87 次
- Neural Basis Models for InterpretabilityFilip Radenovic, Abhimanyu Dubey, Dhruv MahajanNeurIPS 2022 · 被引用 82 次
- Explaining Predictive Uncertainty with Information Theoretic Shapley ValuesDavid S. Watson, Joshua O'Hara, Niek Tax, Richard Mudd 等NeurIPS 2023 · 被引用 56 次
它引用的顶会 Paper4
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana 等CHI 2020 · 被引用 541 次
- Concise Explanations of Neural Networks using Adversarial TrainingPrasad Chalasani, Jiefeng Chen, Amrita Roy Chowdhury, Xi Wu 等ICML 2020 · 被引用 148 次
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay 等ICML 2021 · 被引用 71 次
相关 Paper
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 被引用 56 次
- Shaping Up SHAP: Enhancing Stability through Layer-Wise Neighbor SelectionGwladys Kelodjou, Laurence Rozé, Véronique Masson, Luis Galárraga 等AAAI 2024 · 被引用 24 次
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 被引用 7 次
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 被引用 6 次
- Sparse and Faithful Local Explanations with Piecewise Linear SurrogatesYixin Wang, Yucheng DongICML 2026
