Interpreting Language Models with Contrastive Explanations
Kayo Yin, Graham Neubig
摘要
Model interpretability methods are often used to explain NLP model decisions on tasks such as text classification, where the output space is relatively small. However, when applied to language generation, where the output space often consists of tens of thousands of tokens, these methods are unable to provide informative explanations. Language models must consider various features to predict a token, such as its part of speech, number, tense, or semantics. Existing explanation methods conflate evidence for all these features into a single explanation, which is less interpretable for human understanding. To disentangle the different decisions in language modeling, we focus on explaining language models contrastively: we look for salient input tokens that explain why the model predicted one token instead of another. We demonstrate that contrastive explanations are quantifiably better than non-contrastive explanations in verifying major grammatical phenomena, and that they significantly improve contrastive model simulatability for human observers. We also identify groups of contrastive decisions where the model uses similar evidence, and we are able to characterize what input tokens models use during various language generation decisions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Understanding the Prompt SensitivityYang Liu, Chenhui ChuACL 2026 · 被引用 192 次
- ContextCite: Attributing Model Generation to ContextBenjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander MadryNeurIPS 2024 · 被引用 118 次
- Aligner: Efficient Alignment by Learning to CorrectJiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong 等NeurIPS 2024 · 被引用 115 次
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun 等NeurIPS 2023 · 被引用 87 次
- MambaLRP: Explaining Selective State Space Sequence ModelsFarnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver EberleNeurIPS 2024 · 被引用 44 次
它引用的顶会 Paper4
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar 等EMNLP 2021 · 被引用 12 次
- Do Context-Aware Translation Models Pay the Right Attention?Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary 等ACL 2021
- Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber 等ACL 2021
相关 Paper
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 被引用 9 次
- Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesEleftheria Briakou, Navita Goyal, Marine CarpuatEMNLP 2023
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 被引用 8 次
- KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language InferenceQianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li 等ACL 2021
- Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsSasha Boguraev, Christopher Potts, Kyle MahowaldEMNLP 2025
