S-LIME: Stabilized-LIME for Model Explanation
Zhengze Zhou, Giles Hooker, Fei Wang
摘要
An increasing number of machine learning models have been deployed in domains with high stakes such as finance and healthcare. Despite their superior performances, many models are black boxes in nature which are hard to explain. There are growing efforts for researchers to develop methods to interpret these black-box models. Post hoc explanations based on perturbations, such as LIME [39] , are widely used approaches to interpret a machine learning model after it has been built. This class of methods has been shown to exhibit large instability, posing serious challenges to the effectiveness of the method itself and harming user trust. In this paper, we propose S-LIME, which utilizes a hypothesis testing framework based on central limit theorem for determining the number of perturbation points needed to guarantee stability of the resulting explanation. Experiments on both simulated and real world data sets are provided to demonstrate the effectiveness of our method. CCS CONCEPTS • Computing methodologies → Feature selection; Supervised learning by classification; • Mathematics of computing → Hypothesis testing and confidence interval computation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Falcon: A Privacy-Preserving and Interpretable Vertical Federated Learning SystemYuncheng Wu, Naili Xing, Gang Chen, Tien Tuan Anh Dinh 等VLDB 2023 · 被引用 47 次
- ControlBurn: Feature Selection by Sparse ForestsBrian Liu, Miaolan Xie, Madeleine UdellKDD 2021 · 被引用 6 次
- GiLOT: Interpreting Generative Language Models via Optimal TransportXuhong Li, Jiamin Chen, Yekun Chai, Haoyi XiongICML 2024 · 被引用 6 次
- Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersZhen Li, Ruqian Zhang, Deqing Zou, Ning Wang 等ASE 2023 · 被引用 4 次
- Feature Responsiveness Scores: Model-Agnostic Explanations for RecourseSeung Hyun Cheon, Anneke Wernerfelt, Sorelle A. Friedler, Berk UstunICLR 2025
它引用的顶会 Paper1
相关 Paper
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay 等ICML 2021 · 被引用 71 次
- GLIME: General, Stable and Local LIME ExplanationZeren Tan, Yang Tian, Jian LiNeurIPS 2023 · 被引用 56 次
- "Are Your Explanations Reliable?" Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial AttackChristopher Burger, Lingwei Chen, Thai LeEMNLP 2023 · 被引用 11 次
- Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc ExplanationsTessa Han, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2022 · 被引用 126 次
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei 等KDD 2026
