Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
Hanjie Chen, Yangfeng Ji
摘要
To build an interpretable neural text classifier, most of the prior work has focused on designing inherently interpretable models or finding faithful explanations. A new line of work on improving model interpretability has just started, and many existing methods require either prior information or human annotations as additional inputs in training. To address this limitation, we propose the variational word mask (VMASK) method to automatically learn task-specific important words and reduce irrelevant information on classification, which ultimately improves the interpretability of model predictions. The proposed method is evaluated with three neural text classifiers (CNN, LSTM, and BERT) on seven benchmark text classification datasets. Experiments show the effectiveness of VMASK in improving both model prediction accuracy and interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 被引用 121 次
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu 等ICLR 2021 · 被引用 113 次
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and InterpretationHanjie Chen, Yangfeng JiAAAI 2022 · 被引用 31 次
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An 等NeurIPS 2022 · 被引用 24 次
- Local Explanation of Dialogue Response GenerationYi-Lin Tuan, Connor Pryor, Wenhu Chen, Lise Getoor 等NeurIPS 2021 · 被引用 13 次
它引用的顶会 Paper2
相关 Paper
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 被引用 85 次
- Improving Interpretability via Explicit Word Interaction Graph LayerArshdeep Sekhon, Hanjie Chen, Aman Shrivastava, Zhe Wang 等AAAI 2023 · 被引用 8 次
- SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia TsvetkovEMNLP 2021 · 被引用 39 次
- Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text ClassificationGeorge Chrysostomou, Nikolaos AletrasACL 2021
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 被引用 115 次
