Distribution-Based Feature Attribution for Explaining the Predictions of Any Classifier
Xinpeng Li, Kai Ming Ting
Abstract
The proliferation of complex, black-box AI models has intensified the need for techniques that can explain their decisions. Feature attribution methods have become a popular solution for providing post-hoc explanations, yet the field has historically lacked a formal problem definition. This paper addresses this gap by introducing a formal definition for the problem of feature attribution, which stipulates that explanations be supported by an underlying probability distribution represented by the given dataset. Our analysis reveals that many existing model-agnostic methods fail to meet this criterion, while even those that do often possess other limitations. To overcome these challenges, we propose Distributional Feature Attribution eXplanations (DFAX), a novel, model-agnostic method for feature attribution. DFAX is the first feature attribution method to explain classifier predictions directly based on the data distribution. We show through extensive experiments that DFAX is more effective and efficient than state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c4cd323-c29c-43cf-a321-dcc3cd1a2090Builds on3
- Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionChase Walker, Sumit Kumar Jha, Kenny Chen, Rickard EwetzAAAI 2024 · 25 citations
- Locally Invariant Explanations: Towards Stable and Unidirectional Explanations through Local Invariant LearningAmit Dhurandhar, Karthikeyan Natesan Ramamurthy, Kartik Ahuja, Vijay AryaNeurIPS 2023 · 7 citations
- Iterative Search Attribution for Deep Neural NetworksZhiyu Zhu, Huaming Chen, Xinyi Wang, Jiayu Zhang et al.ICML 2024 · 5 citations
Related papers
- LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningXiuyi Jia, Jinchi Li, Yunan Lu, Weiwei LiICML 2025
- Are Data-Driven Explanations Robust Against Out-of-Distribution Data?Tang Li, Fengchun Qiao, Mengmeng Ma, Xi PengCVPR 2023
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
- GEFA: A General Feature Attribution Framework Using Proxy Gradient EstimationYi Cai, Thibaud Ardoin, Gerhard WunderICML 2025
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
