Trust Regions for Explanations via Black-Box Probabilistic Certification
Amit Dhurandhar, Swagatam Haldar, Dennis Wei, Karthikeyan Natesan Ramamurthy
Abstract
Given the black box nature of machine learning models, a plethora of explainability methods have been developed to decipher the factors behind individual decisions. In this paper, we introduce a novel problem of black box (probabilistic) explanation certification. We ask the question: Given a black box model with only query access, an explanation for an example and a quality metric (viz. fidelity, stability), can we find the largest hypercube (i.e., ball) centered at the example such that when the explanation is applied to all examples within the hypercube, (with high probability) a quality criterion is met (viz. fidelity greater than some value)? Being able to efficiently find such a trust region has multiple benefits: i) insight into model behavior in a region, with a guarantee; ii) ascertained stability of the explanation; iii) explanation reuse, which can save time, energy and money by not having to find explanations for every example; and iv) a possible meta-metric to compare explanation methods. Our contributions include formalizing this problem, proposing solutions, providing theoretical guarantees for these solutions that are computable, and experimentally showing their efficacy on synthetic and real data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a4c5604-7fd7-483d-84f3-34f538b9897dCited by top-tier papers1
Ask how each one uses itBuilds on9
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- On Computing Probabilistic Explanations for Decision TreesMarcelo Arenas, Pablo Barceló, Miguel A. Romero Orth, Bernardo SubercaseauxNeurIPS 2022 · 57 citations
- Consistent Counterfactuals for Deep ModelsEmily Black, Zifan Wang, Matt FredriksonICLR 2022 · 56 citations
- Robust Counterfactual Explanations for Neural Networks With Probabilistic GuaranteesFaisal Hamman, Erfaun Noorani, Saumitra Mishra, Daniele Magazzeni et al.ICML 2023 · 54 citations
- Model Agnostic Multilevel ExplanationsKarthikeyan Natesan Ramamurthy, Bhanukiran Vinzamuri, Yunfeng Zhang, Amit DhurandharNeurIPS 2020 · 48 citations
Related papers
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Is this the Right Neighborhood? Accurate and Query Efficient Model Agnostic ExplanationsAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Karthikeyan ShanmugamNeurIPS 2022 · 9 citations
- Making the Classification Explanation Faithful to the Confidence ScoreJian-Xun Mi, Lu Pan, Weisheng LiCVPR 2026
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei et al.KDD 2026
