Distribution-Based Feature Attribution for Explaining the Predictions of Any Classifier
Xinpeng Li, Kai Ming Ting
摘要
The proliferation of complex, black-box AI models has intensified the need for techniques that can explain their decisions. Feature attribution methods have become a popular solution for providing post-hoc explanations, yet the field has historically lacked a formal problem definition. This paper addresses this gap by introducing a formal definition for the problem of feature attribution, which stipulates that explanations be supported by an underlying probability distribution represented by the given dataset. Our analysis reveals that many existing model-agnostic methods fail to meet this criterion, while even those that do often possess other limitations. To overcome these challenges, we propose Distributional Feature Attribution eXplanations (DFAX), a novel, model-agnostic method for feature attribution. DFAX is the first feature attribution method to explain classifier predictions directly based on the data distribution. We show through extensive experiments that DFAX is more effective and efficient than state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionChase Walker, Sumit Kumar Jha, Kenny Chen, Rickard EwetzAAAI 2024 · 被引用 25 次
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 被引用 7 次
- Iterative Search Attribution for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Xinyi Wang, Jiayu Zhang 等ICML 2024 · 被引用 5 次
相关 Paper
- LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningXiuyi Jia, Jinchi Li, Yunan Lu, Weiwei LiICML 2025
- Are Data-Driven Explanations Robust Against Out-of-Distribution Data?Tang Li, Fengchun Qiao, Mengmeng Ma, Xi PengCVPR 2023
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell 等NeurIPS 2020 · 被引用 83 次
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
