Is this the Right Neighborhood? Accurate and Query Efficient Model Agnostic Explanations
Amit Dhurandhar, Karthikeyan Natesan Ramamurthy, Karthikeyan Shanmugam
摘要
There have been multiple works that try to ascertain explanations for decisions of black box models on particular inputs by perturbing the input or by sampling around it, creating a neighborhood and then fitting a sparse (linear) model (e.g. LIME). Many of these methods are unstable and so more recent work tries to find stable or robust alternatives. However, stable solutions may not accurately represent the behavior of the model around the input. Thus, the question we ask in this paper is are we approximating the local boundary around the input accurately? In particular, are we sampling the right neighborhood so that a linear approximation of the black box is faithful to its true behavior around that input given that the black box can be highly non-linear (viz. deep relu network with many linear pieces). It is difficult to know the correct neighborhood width (or radius) as too small a width can lead to a bad condition number of the inverse covariance matrix of function fitting procedures resulting in unstable predictions, while too large a width may lead to accounting for multiple linear pieces and consequently a poor local approximation. We in this paper propose a simple approach that is robust across neighborhood widths in recovering faithful local explanations. In addition to a naive implementation of our approach which can still be accurate, we propose a novel adaptive neighborhood sampling scheme (ANS) that we formally show can be much more sample and query efficient. We then empirically evaluate our approach on real data where our explanations are significantly more sample and query efficient than the competitors, while also being faithful and stable across different widths.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Improving ML-based Binary Function Similarity Detection by Assessing and Deprioritizing Control Flow Graph FeaturesJialai Wang, Chao Zhang, Longfei Chen, Yi Rong 等USENIX Security 2024 · 被引用 15 次
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 被引用 7 次
- Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal ExplanationXuexin Chen, Ruichu Cai, Zhengting Huang, Yuxuan Zhu 等ICML 2024 · 被引用 5 次
- Trust Regions for Explanations via Black-Box Probabilistic CertificationAmit Dhurandhar, Swagatam Haldar, Dennis Wei, Karthikeyan Natesan RamamurthyICML 2024 · 被引用 3 次
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
相关 Paper
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei 等KDD 2026
- Sparse and Faithful Local Explanations with Piecewise Linear SurrogatesYixin Wang, Yucheng DongICML 2026
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 被引用 56 次
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay 等ICML 2021 · 被引用 71 次
- S-LIME: Stabilized-LIME for Model ExplanationZhengze Zhou, Giles Hooker, Fei WangKDD 2021 · 被引用 98 次
