A Unifying Framework to the Analysis of Interaction Methods using Synergy Functions
Daniel Lundström, Meisam Razaviyayn
Abstract
Deep learning has revolutionized many areas of machine learning, from computer vision to natural language processing, but these high-performance models are generally "black box." Explaining such models would improve transparency and trust in AI-powered decision making and is necessary for understanding other practical needs such as robustness and fairness. A popular means of enhancing model transparency is to quantify how individual inputs contribute to model outputs (called attributions) and the magnitude of interactions between groups of inputs. A growing number of these methods import concepts and results from game theory to produce attributions and interactions. This work presents a unifying framework for game-theory-inspired attribution and k th -order interaction methods. We show that, given modest assumptions, a unique full account of interactions between features, called synergies, is possible in the continuous input setting. We identify how various methods are characterized by their policy of distributing synergies. We establish that gradientbased methods are characterized by their actions on monomials, a type of synergy function, and introduce unique gradient-based methods. We show that the combination of various criteria uniquely defines the attribution/interaction methods. Thus, the community needs to identify goals and contexts when developing and employing attribution and interaction methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61fdcd47-ef67-4e78-8c40-500041920006Cited by top-tier papers1
Ask how each one uses itBuilds on12
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 109 citations
- A Rigorous Study of Integrated Gradients Method and Extensions to Internal Neuron AttributionsDaniel Lundström, Tianjian Huang, Meisam RazaviyaynICML 2022 · 85 citations
Related papers
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
- A Unified Taylor Framework for Revisiting Attribution MethodsHuiqi Deng, Na Zou, Mengnan Du, Weifu Chen et al.AAAI 2021 · 25 citations
- Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsSandipan Sikdar, Parantapa Bhattacharya, Kieran HeeseACL 2021
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 27 citations
- Training Characteristic Functions with Reinforcement Learning: XAI-methods play Connect FourStephan Wäldchen, Sebastian Pokutta, Felix HuberICML 2022 · 9 citations
