The Many Shapley Values for Model Explanation
Mukund Sundararajan, Amir Najmi
Abstract
The Shapley value has become a popular method to attribute the prediction of a machine-learning model on an input to its base features. The use of the Shapley value is justified by citing [16] showing that it is the unique method that satisfies certain good properties (axioms). There are, however, a multiplicity of ways in which the Shapley value is operationalized in the attribution problem. These differ in how they reference the model, the training data, and the explanation context. These give very different results, rendering the uniqueness result meaningless. Furthermore, we find that previously proposed approaches can produce counterintuitive attributions in theory and in practice---for instance, they can assign non-zero attributions to features that are not even referenced by the model. In this paper, we use the axiomatic approach to study the differences between some of the many operationalizations of the Shapley value for attribution, and propose a technique called Baseline Shapley (BShap) that is backed by a proper uniqueness result. We also contrast BShap with Integrated Gradients, another extension of Shapley value to the continuous setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28009a34-a998-4002-bfb9-0d80f993e7eeCited by top-tier papers99
- On the Tractability of SHAP ExplanationsGuy Van den Broeck, Anton Lykov, Maximilian Schleich, Dan SuciuAAAI 2021 · 485 citations
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
- Shapley explainability on the data manifoldChristopher Frye, Damien de Mijolla, Tom Begley, Laurence Cowton et al.ICLR 2021 · 125 citations
- Additive MIL: Intrinsically Interpretable Multiple Instance Learning for PathologySyed Ashar Javed, Dinkar Juyal, Harshith Padigela, Amaro Taylor-Weiner et al.NeurIPS 2022 · 124 citations
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 121 citations
Builds on1
Related papers
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 60 citations
- RankSHAP: Shapley Value Based Feature Attributions for Learning to RankTanya Chowdhury, Yair Zick, James AllanICLR 2025
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- Joint Shapley values: a measure of joint feature importanceChris Harris, Richard Pymar, Colin RowatICLR 2022 · 32 citations
- Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex ModelsTom Heskes, Evi Sijben, Ioan Gabriel Bucur, Tom ClaassenNeurIPS 2020 · 235 citations
