Trust Regions for Explanations via Black-Box Probabilistic Certification
Amit Dhurandhar, Swagatam Haldar, Dennis Wei, Karthikeyan Natesan Ramamurthy
摘要
Given the black box nature of machine learning models, a plethora of explainability methods have been developed to decipher the factors behind individual decisions. In this paper, we introduce a novel problem of black box (probabilistic) explanation certification. We ask the question: Given a black box model with only query access, an explanation for an example and a quality metric (viz. fidelity, stability), can we find the largest hypercube (i.e., ball) centered at the example such that when the explanation is applied to all examples within the hypercube, (with high probability) a quality criterion is met (viz. fidelity greater than some value)? Being able to efficiently find such a trust region has multiple benefits: i) insight into model behavior in a region, with a guarantee; ii) ascertained stability of the explanation; iii) explanation reuse, which can save time, energy and money by not having to find explanations for every example; and iv) a possible meta-metric to compare explanation methods. Our contributions include formalizing this problem, proposing solutions, providing theoretical guarantees for these solutions that are computable, and experimentally showing their efficacy on synthetic and real data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov 等S&P 2018 · 被引用 987 次
- On Computing Probabilistic Explanations for Decision TreesMarcelo Arenas, Pablo Barceló, Miguel A. Romero Orth, Bernardo SubercaseauxNeurIPS 2022 · 被引用 57 次
- Consistent Counterfactuals for Deep ModelsEmily Black, Zifan Wang, Matt FredriksonICLR 2022 · 被引用 56 次
- Robust Counterfactual Explanations for Neural Networks With Probabilistic GuaranteesFaisal Hamman, Erfaun Noorani, Saumitra Mishra, Daniele Magazzeni 等ICML 2023 · 被引用 54 次
- Model Agnostic Multilevel ExplanationsKarthikeyan Natesan Ramamurthy, Bhanukiran Vinzamuri, Yunfeng Zhang, Amit DhurandharNeurIPS 2020 · 被引用 48 次
相关 Paper
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- Is this the Right Neighborhood? Accurate and Query Efficient Model Agnostic ExplanationsAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Karthikeyan ShanmugamNeurIPS 2022 · 被引用 9 次
- Making the Classification Explanation Faithful to the Confidence ScoreJian-Xun Mi, Lu Pan, Weisheng LiCVPR 2026
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 被引用 35 次
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei 等KDD 2026
