Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP Models
Sandipan Sikdar, Parantapa Bhattacharya, Kieran Heese
Abstract
In this paper, we introduce Integrated Directional Gradients (IDG), a method for attributing importance scores to groups of features, indicating their relevance to the output of a neural network model for a given input. The success of Deep Neural Networks has been attributed to their ability to capture higher level feature interactions. Hence, in the last few years capturing the importance of these feature interactions has received increased prominence in ML interpretability literature. In this paper, we formally define the feature group attribution problem and outline a set of axioms that any intuitive feature group attribution method should satisfy. Earlier, cooperative game theory inspired axiomatic methods only borrowed axioms from solution concepts (such as Shapley value) for individual feature attributions and introduced their own extensions to model interactions. In contrast, our formulation is inspired by axioms satisfied by characteristic functions as well as solution concepts in cooperative game theory literature. We believe that characteristic functions are much better suited to model importance of groups compared to just solution concepts. We demonstrate that our proposed method, IDG, satisfies all the axioms. Using IDG we analyze two state-of-the-art text classifiers on three benchmark datasets for sentiment analysis. Our experiments show that IDG is able to effectively capture semantic interactions in linguistic models via negations and conjunctions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ad438dc-7f1c-4a46-bb4d-490ab82fdebfCited by top-tier papers13
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm et al.EMNLP 2022 · 29 citations
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language ModelsYihao Zhang, Zeming Wei, Jun Sun, Meng SunNeurIPS 2024 · 16 citations
- AD-KD: Attribution-Driven Knowledge Distillation for Language Model CompressionSiyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang et al.ACL 2023 · 10 citations
- Explaining Interactions Between Text SpansSagnik Ray Choudhury, Pepa Atanasova, Isabelle AugensteinEMNLP 2023 · 4 citations
- A Unifying Framework to the Analysis of Interaction Methods using Synergy FunctionsDaniel Lundström, Meisam RazaviyaynICML 2023 · 4 citations
Builds on10
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 799 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 147 citations
Related papers
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
- A Rigorous Study of Integrated Gradients Method and Extensions to Internal Neuron AttributionsDaniel Lundström, Tianjian Huang, Meisam RazaviyaynICML 2022 · 85 citations
- GStarX: Explaining Graph Neural Networks with Structure-Aware Cooperative GamesShichang Zhang, Yozen Liu, Neil Shah, Yizhou SunNeurIPS 2022 · 79 citations
- H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image ClassifiersAyushi Mehrotra, Dipkamal Bhusal, Michael Clifford, Nidhi RastogiCVPR 2026 · 1 citation
- Discretized Integrated Gradients for Explaining Language ModelsSoumya Sanyal, Xiang RenEMNLP 2021 · 34 citations
