Theoretical Behavior of XAI Methods in the Presence of Suppressor Variables
Rick Wilming, Leo Kieslich, Benedict Clark, Stefan Haufe
摘要
In recent years, the community of 'explainable artificial intelligence' (XAI) has created a vast body of methods to bridge a perceived gap between model 'complexity' and 'interpretability'. However, a concrete problem to be solved by XAI methods has not yet been formally stated. As a result, XAI methods are lacking theoretical and empirical evidence for the 'correctness' of their explanations, limiting their potential use for quality-control and transparency purposes. At the same time, Haufe et al. ( 2014 ) showed, using simple toy examples, that even standard interpretations of linear models can be highly misleading. Specifically, high importance may be attributed to so-called suppressor variables lacking any statistical relation to the prediction target. This behavior has been confirmed empirically for a large array of XAI methods in Wilming et al. (2022) . Here, we go one step further by deriving analytical expressions for the behavior of a variety of popular XAI methods on a simple two-dimensional binary classification problem involving Gaussian class-conditional distributions. We show that the majority of the studied approaches will attribute non-zero importance to a non-class-related suppressor feature in the presence of correlated noise. This poses important limitations on the interpretations and conclusions that the outputs of these XAI methods can afford.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Minimizing False-Positive Attributions in Explanations of Non-Linear ModelsAnders Gjølbye, Stefan Haufe, Lars Kai HansenNeurIPS 2025 · 被引用 3 次
- Correcting misinterpretations of additive modelsBenedict Clark, Rick Wilming, Hjalmar Schulz, Rustam Zhumagambetov 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- A Causality Inspired Framework for Model InterpretationChenwang Wu, Xiting Wang, Defu Lian, Xing Xie 等KDD 2023 · 被引用 22 次
- (Mis)Communicating with our AI SystemsLaura Cros Vila, Bob L. T. SturmCHI 2025 · 被引用 2 次
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- A Psychological Theory of ExplainabilityScott Cheng-Hsin Yang, Tomas Folke, Patrick ShaftoICML 2022 · 被引用 21 次
