Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Hanjie Chen, Guangtao Zheng, Yangfeng Ji
摘要
Generating explanations for neural networks has become crucial for their applications in real-world with respect to reliability and trustworthiness. In natural language processing, existing methods usually provide important features which are words or phrases selected from an input text as an explanation, but ignore the interactions between them. It poses challenges for humans to interpret an explanation and connect it to model prediction. In this work, we build hierarchical explanations by detecting feature interactions. Such explanations visualize how words and phrases are combined at different levels of the hierarchy, which can help users understand the decision-making of blackbox models. The proposed method is evaluated with three neural text classifiers (LSTM, CNN, and BERT) on two benchmark datasets, via both automatic and human evaluations. Experiments show the effectiveness of the proposed method in providing explanations that are both faithful to models and interpretable to humans.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 被引用 282 次
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu 等ICLR 2021 · 被引用 113 次
- AttCAT: Explaining Transformers via Attentive Class Activation TokensYao Qiang, Deng Pan, Chengyin Li, Xin Li 等NeurIPS 2022 · 被引用 66 次
- Building Interpretable Interaction Trees for Deep NLP ModelsDie Zhang, Hao Zhang, Huilin Zhou, Xiaoyi Bao 等AAAI 2021 · 被引用 43 次
- : Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language ModelsXingbo Wang, Renfei Huang, Zhihua Jin, Tianqing Fang 等IEEE VIS 2023 · 被引用 18 次
它引用的顶会 Paper3
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 被引用 774 次
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue 等ICLR 2020 · 被引用 55 次
- LS-Tree: Model Interpretation When the Data Are LinguisticJianbo Chen, Michael I. JordanAAAI 2020 · 被引用 19 次
相关 Paper
- Learning Variational Word Masks to Improve the Interpretability of Neural Text ClassifiersHanjie Chen, Yangfeng JiEMNLP 2020 · 被引用 45 次
- Concept Bottleneck Large Language ModelsChung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei WengICLR 2025
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia TsvetkovEMNLP 2021 · 被引用 39 次
- Let the CAT out of the bag: Contrastive Attributed explanations for TextSaneem A. Chemmengath, Amar Prakash Azad, Ronny Luss, Amit DhurandharEMNLP 2022 · 被引用 6 次
