Building Interpretable Interaction Trees for Deep NLP Models
Die Zhang, Hao Zhang, Huilin Zhou, Xiaoyi Bao, Da Huo, Ruizhao Chen, Xu Cheng, Mengyue Wu, Quanshi Zhang
Abstract
This paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituents in a sentence. The interaction is defined based on Shapley values of words, which are considered as an unbiased estimation of word contributions to the network prediction. Our method is used to quantify word interactions encoded inside the BERT, ELMo, LSTM, CNN, and Transformer networks. Experimental results have provided a new perspective to understand these DNNs, and have demonstrated the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1308e8a-a2b4-4659-857c-c8713b9584e5Cited by top-tier papers15
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 73 citations
- Interpreting and Boosting Dropout from a Game-Theoretic ViewHao Zhang, Sen Li, Yinchao Ma, Mingjie Li et al.ICLR 2021 · 53 citations
- Does a Neural Network Really Encode Symbolic Concepts?Mingjie Li, Quanshi ZhangICML 2023 · 35 citations
- Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesDongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren et al.NeurIPS 2023 · 28 citations
Builds on5
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
- Knowledge Consistency between Neural Networks and BeyondRuofan Liang, Tianlin Li, Longfei Li, Jing Wang et al.ICLR 2020 · 30 citations
- LS-Tree: Model Interpretation When the Data Are LinguisticJianbo Chen, Michael I. JordanAAAI 2020 · 19 citations
- Explaining Knowledge Distillation by Quantifying the KnowledgeXu Cheng, Zhefan Rao, Yilan Chen, Quanshi ZhangCVPR 2020
Related papers
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang et al.AAAI 2021 · 70 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- Using Shapley interactions to understand how models use structureDivyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain et al.ACL 2025 · 1 citation
- HarsanyiNet: Computing Accurate Shapley Values in a Single Forward PropagationLu Chen, Siyu Lou, Keyan Zhang, Jin Huang et al.ICML 2023 · 17 citations
- Interpreting Attributions and Interactions of Adversarial AttacksXin Wang, Shuyun Lin, Hao Zhang, Yufei Zhu et al.ICCV 2021 · 20 citations
