Rethinking Shapley Value for Negative Interactions in Non-convex Games
Wonjoon Chang, Myeongjin Lee, Jaesik Choi
Abstract
We study causal interactions for payoff allocation in cooperative game theory, including quantifying feature attribution for deep learning models. Most feature attribution methods mainly stem from the criteria of the Shapley value, which assigns fair payoffs to players based on their expected contribution in a cooperative game. However, interactions between players in the game do not explicitly appear in the original formulation of the Shapley value. In this work, we reformulate the Shapley value to clarify the role of interactions and discuss implicit assumptions from a game-theoretical perspective. Our theoretical analysis demonstrates that when negative interactions exist—common in deep learning models—the efficiency axiom can lead to the undervaluation of attributions or payoffs. We suggest a new allocation rule that decomposes contributions into interactions and aggregates positive parts for non-convex games. Furthermore, we propose an approximation algorithm to reduce the cost of interaction computation which can be applied to differentiable functions such as deep learning models. Our approach mitigates counterintuitive attribution outcomes observed in existing methods, ensuring that features critical to a model’s decision receive appropriate attribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksWoo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf et al.AAAI 2020 · 109 citations
- Shapley Residuals: Quantifying the limits of the Shapley value for explanationsIndra Kumar, Carlos Scheidegger, Suresh Venkatasubramanian, Sorelle A. FriedlerNeurIPS 2021 · 82 citations
- SHAP-IQ: Unified Approximation of any-order Shapley InteractionsFabian Fumagalli, Maximilian Muschalik, Patrick Kolpaczki, Eyke Hüllermeier et al.NeurIPS 2023 · 80 citations
Related papers
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- You Mostly Walk Alone: Analyzing Feature Attribution in Trajectory PredictionOsama Makansi, Julius von Kügelgen, Francesco Locatello, Peter Vincent Gehler et al.ICLR 2022 · 35 citations
- Neural Payoff Machines: Predicting Fair and Stable Payoff Allocations Among Team MembersDaphne Cornelisse, Thomas Rood, Yoram Bachrach, Mateusz Malinowski et al.NeurIPS 2022 · 10 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
- Unlocking the Game: Estimating Games in Möbius Representation for Explanation and High-Order Interaction DetectionMajid Mohammadi, Ilaria Tiddi, Annette ten TeijeAAAI 2025 · 4 citations
