Interpreting Language Models with Contrastive Explanations
Kayo Yin, Graham Neubig
Abstract
Model interpretability methods are often used to explain NLP model decisions on tasks such as text classification, where the output space is relatively small. However, when applied to language generation, where the output space often consists of tens of thousands of tokens, these methods are unable to provide informative explanations. Language models must consider various features to predict a token, such as its part of speech, number, tense, or semantics. Existing explanation methods conflate evidence for all these features into a single explanation, which is less interpretable for human understanding. To disentangle the different decisions in language modeling, we focus on explaining language models contrastively: we look for salient input tokens that explain why the model predicted one token instead of another. We demonstrate that contrastive explanations are quantifiably better than non-contrastive explanations in verifying major grammatical phenomena, and that they significantly improve contrastive model simulatability for human observers. We also identify groups of contrastive decisions where the model uses similar evidence, and we are able to characterize what input tokens models use during various language generation decisions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05801638-e18d-4369-a8fc-835f0abd384cCited by top-tier papers27
- Understanding the Prompt SensitivityYang Liu, Chenhui ChuACL 2026 · 192 citations
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 118 citations
- Aligner: Efficient Alignment by Learning to CorrectJiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong et al.NeurIPS 2024 · 115 citations
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun et al.NeurIPS 2023 · 87 citations
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 44 citations
Builds on4
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar et al.EMNLP 2021 · 12 citations
- Do Context-Aware Translation Models Pay the Right Attention?Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary et al.ACL 2021
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber et al.ACL 2021
Related papers
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 9 citations
- Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesEleftheria Briakou, Navita Goyal, Marine CarpuatEMNLP 2023
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 8 citations
- KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language InferenceQianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li et al.ACL 2021
- Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsSasha Boguraev, Christopher Potts, Kyle MahowaldEMNLP 2025
