Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
Hanjie Chen, Yangfeng Ji
Abstract
To build an interpretable neural text classifier, most of the prior work has focused on designing inherently interpretable models or finding faithful explanations. A new line of work on improving model interpretability has just started, and many existing methods require either prior information or human annotations as additional inputs in training. To address this limitation, we propose the variational word mask (VMASK) method to automatically learn task-specific important words and reduce irrelevant information on classification, which ultimately improves the interpretability of model predictions. The proposed method is evaluated with three neural text classifiers (CNN, LSTM, and BERT) on seven benchmark text classification datasets. Experiments show the effectiveness of VMASK in improving both model prediction accuracy and interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87429090-d0b7-4224-8fcf-c162ac204f83Cited by top-tier papers19
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 121 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and InterpretationHanjie Chen, Yangfeng JiAAAI 2022 · 31 citations
- DARE: Disentanglement-Augmented Rationale ExtractionLinan Yue, Qi Liu, Yichao Du, Yanqing An et al.NeurIPS 2022 · 24 citations
- Local Explanation of Dialogue Response GenerationYi-Lin Tuan, Connor Pryor, Wenhu Chen, Lise Getoor et al.NeurIPS 2021 · 13 citations
Builds on2
- Restricting the Flow: Information Bottlenecks for AttributionKarl Schulz, Leon Sixt, Federico Tombari, Tim LandgrafICLR 2020 · 220 citations
- Regularizing Black-box Models for Improved InterpretabilityGregory Plumb, Maruan Al-Shedivat, Ángel Alexander Cabrera, Adam Perer et al.NeurIPS 2020 · 90 citations
Related papers
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 85 citations
- Improving Interpretability via Explicit Word Interaction Graph LayerArshdeep Sekhon, Hanjie Chen, Aman Shrivastava, Zhe Wang et al.AAAI 2023 · 8 citations
- SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia TsvetkovEMNLP 2021 · 39 citations
- Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text ClassificationGeorge Chrysostomou, Nikolaos AletrasACL 2021
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 115 citations
