NILE : Natural Language Inference with Faithful Natural Language Explanations
Sawan Kumar, Partha P. Talukdar
摘要
The recent growth in the popularity and success of deep learning models on NLP classification tasks has accompanied the need for generating some form of natural language explanation of the predicted labels. Such generated natural language (NL) explanations are expected to be faithful, i.e., they should correlate well with the model's internal decision making. In this work, we focus on the task of natural language inference (NLI) and address the following question: can we build NLI systems which produce labels with high accuracy, while also generating faithful explanations of its decisions? We propose Naturallanguage Inference over Label-specific Explanations (NILE), a novel NLI method which utilizes auto-generated label-specific NL explanations to produce labels along with its faithful explanation. We demonstrate NILE's effectiveness over previously reported methods through automated and human evaluation of the produced labels and explanations. Our evaluation of NILE also supports the claim that accurate systems capable of providing testable explanations of their decisions can be designed. We discuss the faithfulness of NILE's explanations in terms of sensitivity of the decisions to the corresponding explanations. We argue that explicit evaluation of faithfulness, in addition to label and explanation accuracy, is an important step in evaluating model's explanations. Further, we demonstrate that task-specific probes are necessary to establish such sensitivity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language TasksMaxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde 等ICCV 2021 · 被引用 115 次
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 等ICML 2024 · 被引用 90 次
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 被引用 52 次
- UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionAaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 等ICML 2022 · 被引用 48 次
- Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionFaeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin ChoiAAAI 2021 · 被引用 46 次
它引用的顶会 Paper3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
相关 Paper
- Logical Satisfiability of Counterfactuals for Faithful Explanations in NLISuzanna Sia, Anton Belyy, Amjad Almahairi, Madian Khabsa 等AAAI 2023 · 被引用 17 次
- Graph-Guided Textual Explanation Generation FrameworkShuzhou Yuan, Jingyi Sun, Ran Zhang, Michael Färber 等EMNLP 2025 · 被引用 1 次
- Zero-Shot Natural Language ExplanationsFawaz Sammani, Nikos DeligiannisICLR 2025
- MaNtLE: Model-agnostic Natural Language ExplainerRakesh R. Menon, Kerem Zaman, Shashank SrivastavaEMNLP 2023 · 被引用 1 次
- Alignment Rationale for Natural Language InferenceZhongtao Jiang, Yuanzhe Zhang, Zhao Yang, Jun Zhao 等ACL 2021
