Measuring Association Between Labels and Free-Text Rationales
Sarah Wiegreffe, Ana Marasovic, Noah A. Smith
摘要
In interpretable NLP, we require faithful rationales that reflect the model's decision-making process for an explained instance. While prior work focuses on extractive rationales (a subset of the input words), we investigate their less-studied counterpart: free-text natural language rationales. We demonstrate that pipelines, existing models for faithful extractive rationalization on information-extraction style tasks, do not extend as reliably to "reasoning" tasks requiring free-text rationales. We turn to models that jointly predict and rationalize, a class of widely used high-performance models for free-text rationalization whose faithfulness is not yet established. We define label-rationale association as a necessary property for faithfulness: the internal mechanisms of the model producing the label and the rationale must be meaningfully correlated. We propose two measurements to test this property: robustness equivalence and feature importance agreement. We find that state-of-the-art T5-based joint models exhibit both properties for rationalizing commonsense question-answering and natural language inference, indicating their potential for producing faithful free-text rationales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Evaluation Guided Beam Search for ReasoningYuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao 等NeurIPS 2023 · 被引用 316 次
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 被引用 272 次
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
它引用的顶会 Paper12
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
- SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersDheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia TsvetkovEMNLP 2021 · 被引用 39 次
相关 Paper
- QUASER: Question Answering with Scalable Extractive RationalizationAsish Ghoshal, Srinivasan Iyer, Bhargavi Paranjape, Kushal Lakhotia 等SIGIR 2022 · 被引用 2 次
- RORA: Robust Free-Text Rationale EvaluationZhengping Jiang, Yining Lu, Hanjie Chen, Daniel Khashabi 等ACL 2024
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 被引用 40 次
- Does Self-Rationalization Improve Robustness to Spurious Correlations?Alexis Ross, Matthew E. Peters, Ana MarasovicEMNLP 2022 · 被引用 4 次
- Towards Trustworthy Explanation: On Causal RationalizationWenbo Zhang, Tong Wu, Yunlong Wang, Yong Cai 等ICML 2023 · 被引用 25 次
