Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, Xiang Ren
Abstract
The impressive performance of neural networks on natural language processing tasks attributes to their ability to model complicated word and phrase compositions. To explain how the model handles semantic compositions, we study hierarchical explanation of neural network predictions. We identify non-additivity and context independent importance attributions within hierarchies as two desirable properties for highlighting word and phrase compositions. We show some prior efforts on hierarchical explanations, e.g. contextual decomposition, do not satisfy the desired properties mathematically, leading to inconsistent explanation quality in different models. In this paper, we start by proposing a formal and general way to quantify the importance of each word and phrase. Following the formulation, we propose Sampling and Contextual Decomposition (SCD) algorithm and Sampling and Occlusion (SOC) algorithm. Human and metrics evaluation on both LSTM models and BERT Transformer models on multiple datasets show that our algorithms outperform prior hierarchical explanation algorithms. Our algorithms help to visualize semantic composition captured by models, extract classification rules and improve human trust of models. Project page: this https URL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce9c76fd-c695-450f-bcf9-fd3ee373bc94Cited by top-tier papers27
- Interpreting Graph Neural Networks for NLP With Differentiable Edge MaskingMichael Sejr Schlichtkrull, Nicola De Cao, Ivan TitovICLR 2021 · 287 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 109 citations
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
Related papers
- Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsByung-Doh Oh, William SchulerACL 2023 · 1 citation
- Learning Variational Word Masks to Improve the Interpretability of Neural Text ClassifiersHanjie Chen, Yangfeng JiEMNLP 2020 · 45 citations
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri et al.EMNLP 2024 · 3 citations
- Building Interpretable Interaction Trees for Deep NLP ModelsDie Zhang, Hao Zhang, Huilin Zhou, Xiaoyi Bao et al.AAAI 2021 · 43 citations
- SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia TsvetkovEMNLP 2021 · 39 citations
