Contrastive Explanations for Model Interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, Yoav Goldberg
摘要
Contrastive explanations clarify why an event occurred in contrast to another. They are inherently intuitive to humans to both produce and comprehend. We propose a method to produce contrastive explanations in the latent space, via a projection of the input representation, such that only the features that differentiate two potential decisions are captured. Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions. Additionally, for a given input feature, our contrastive explanations can answer for which label, and against which alternative label, is the feature useful. We produce contrastive explanations via both highlevel abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks. Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- Maieutic Prompting: Logically Consistent Reasoning with Recursive ExplanationsJaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman 等EMNLP 2022 · 被引用 72 次
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 被引用 46 次
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 被引用 32 次
- Contrastive Explanations That Anticipate Human Misconceptions Can Improve Human Decision-Making SkillsZana Buçinca, Siddharth Swaroop, Amanda E. Paluch, Finale Doshi-Velez 等CHI 2025 · 被引用 31 次
它引用的顶会 Paper7
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 被引用 458 次
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 被引用 91 次
- Interpretation of NLP models through input marginalizationSiwon Kim, Jihun Yi, Eunji Kim, Sungroh YoonEMNLP 2020 · 被引用 41 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
相关 Paper
- Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesEleftheria Briakou, Navita Goyal, Marine CarpuatEMNLP 2023
- Leveraging Latent Features for Local ExplanationsRonny Luss, Pin-Yu Chen, Amit Dhurandhar, Prasanna Sattigeri 等KDD 2021 · 被引用 11 次
- Rather a Nurse than a Physician - Contrastive Explanations under InvestigationOliver Eberle, Ilias Chalkidis, Laura Cabello, Stephanie BrandlEMNLP 2023 · 被引用 3 次
- KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language InferenceQianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li 等ACL 2021
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri 等EMNLP 2024 · 被引用 3 次
