Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP Models
Sandipan Sikdar, Parantapa Bhattacharya, Kieran Heese
摘要
In this paper, we introduce Integrated Directional Gradients (IDG), a method for attributing importance scores to groups of features, indicating their relevance to the output of a neural network model for a given input. The success of Deep Neural Networks has been attributed to their ability to capture higher level feature interactions. Hence, in the last few years capturing the importance of these feature interactions has received increased prominence in ML interpretability literature. In this paper, we formally define the feature group attribution problem and outline a set of axioms that any intuitive feature group attribution method should satisfy. Earlier, cooperative game theory inspired axiomatic methods only borrowed axioms from solution concepts (such as Shapley value) for individual feature attributions and introduced their own extensions to model interactions. In contrast, our formulation is inspired by axioms satisfied by characteristic functions as well as solution concepts in cooperative game theory literature. We believe that characteristic functions are much better suited to model importance of groups compared to just solution concepts. We demonstrate that our proposed method, IDG, satisfies all the axioms. Using IDG we analyze two state-of-the-art text classifiers on three benchmark datasets for sentiment analysis. Our experiments show that IDG is able to effectively capture semantic interactions in linguistic models via negations and conjunctions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm 等EMNLP 2022 · 被引用 29 次
- Adversarial Representation Engineering: A General Model Editing Framework for Large Language ModelsYihao Zhang, Zeming Wei, Jun Sun, Meng SunNeurIPS 2024 · 被引用 16 次
- AD-KD: Attribution-Driven Knowledge Distillation for Language Model CompressionSiyue Wu, Hongzhan Chen, Xiaojun Quan, Qifan Wang 等ACL 2023 · 被引用 10 次
- Explaining Interactions Between Text SpansSagnik Ray Choudhury, Pepa Atanasova, Isabelle AugensteinEMNLP 2023 · 被引用 4 次
- A Unifying Framework to the Analysis of Interaction Methods using Synergy FunctionsDaniel Lundström, Meisam RazaviyaynICML 2023 · 被引用 4 次
它引用的顶会 Paper10
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 被引用 799 次
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 被引用 458 次
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 被引用 246 次
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 被引用 199 次
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 被引用 147 次
相关 Paper
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
- A Rigorous Study of Integrated Gradients Method and Extensions to Internal Neuron AttributionsDaniel Lundström, Tianjian Huang, Meisam RazaviyaynICML 2022 · 被引用 85 次
- GStarX: Explaining Graph Neural Networks with Structure-Aware Cooperative GamesShichang Zhang, Yozen Liu, Neil Shah, Yizhou SunNeurIPS 2022 · 被引用 79 次
- H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image ClassifiersAyushi Mehrotra, Dipkamal Bhusal, Michael Clifford, Nidhi RastogiCVPR 2026 · 被引用 1 次
- Discretized Integrated Gradients for Explaining Language ModelsSoumya Sanyal, Xiang RenEMNLP 2021 · 被引用 34 次
