Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, Sungroh Yoon
Abstract
To demystify the "black box" property of deep neural networks for natural language processing (NLP), several methods have been proposed to interpret their predictions by measuring the change in prediction probability after erasing each token of an input. Since existing methods replace each token with a predefined value (i.e., zero), the resulting sentence lies out of the training data distribution, yielding misleading interpretations. In this study, we raise the out-of-distribution problem induced by the existing interpretation methods and present a remedy; we propose to marginalize each token out. We interpret various NLP models trained for sentiment analysis and natural language inference using the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf493854-8078-474a-aeec-be7d0df67905Cited by top-tier papers18
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 121 citations
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 32 citations
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm et al.EMNLP 2022 · 29 citations
- Learning to Scaffold: Optimizing Model Explanations for TeachingPatrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins et al.NeurIPS 2022 · 26 citations
- Jointly Attacking Graph Neural Network and its ExplanationsWenqi Fan, Han Xu, Wei Jin, Xiaorui Liu et al.ICDE 2023 · 23 citations
Builds on2
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence ModelsXisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue et al.ICLR 2020 · 55 citations
Related papers
- Faithfulness Measurable Masked Language ModelsAndreas Madsen, Siva Reddy, Sarath ChandarICML 2024 · 6 citations
- Explainability as statistical inferenceHugo Henri Joseph Senetaire, Damien Garreau, Jes Frellsen, Pierre-Alexandre MatteiICML 2023 · 4 citations
- You should evaluate your language model on marginal likelihood over tokenisationsKris Cao, Laura RimellEMNLP 2021 · 6 citations
- DOCTOR: A Simple Method for Detecting Misclassification ErrorsFederica Granese, Marco Romanelli, Daniele Gorla, Catuscia Palamidessi et al.NeurIPS 2021 · 65 citations
- Statistical Hypothesis Testing for Auditing Robustness in Language ModelsPaulius Rauba, Qiyao Wei, Mihaela van der SchaarICML 2025
