Diagnostics-Guided Explanation Generation
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
Abstract
Explanations shed light on a machine learning model's rationales and can aid in identifying deficiencies in its reasoning process. Explanation generation models are typically trained in a supervised way given human explanations. When such annotations are not available, explanations are often selected as those portions of the input that maximise a downstream task's performance, which corresponds to optimising an explanation's Faithfulness to a given model. Faithfulness is one of several so-called diagnostic properties, which prior work has identified as useful for gauging the quality of an explanation without requiring annotations. Other diagnostic properties are Data Consistency, which measures how similar explanations are for similar input instances, and Confidence Indication, which shows whether the explanation reflects the confidence of the model. In this work, we show how to directly optimise for these diagnostic properties when training a model to generate sentence-level explanations, which markedly improves explanation quality, agreement with human rationales, and downstream task performance on three complex reasoning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df7fab7b-822a-40ea-bc44-1b6b0cd07857Cited by top-tier papers6
- Exploring Faithful Rationale for Multi-Hop Fact Verification via Salience-Aware Graph LearningJiasheng Si, Yingjie Zhu, Deyu ZhouAAAI 2023 · 27 citations
- Hop, Union, Generate: Explainable Multi-hop Reasoning without Rationale SupervisionWenting Zhao, Justin T. Chiu, Claire Cardie, Alexander M. RushEMNLP 2023 · 4 citations
- Unsupervised Selective Rationalization with Noise InjectionAdam Storek, Melanie Subbiah, Kathleen R. McKeownACL 2023 · 4 citations
- Automated Justification Production for Claim Veracity in Fact Checking: A Survey on Architectures and ApproachesIslam Eldifrawi, Shengrui Wang, Amine TrabelsiACL 2024 · 4 citations
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?Yifan Wang, Mayank Jobanputra, Ji-Ung Lee, Soyoung Oh et al.ICLR 2026 · 3 citations
Builds on7
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 130 citations
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song et al.ICLR 2020 · 120 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 15 citations
Related papers
- Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic InferenceJiazheng Li, Hanqi Yan, Yulan HeACL 2025
- Graph-Guided Textual Explanation Generation FrameworkShuzhou Yuan, Jingyi Sun, Ran Zhang, Michael Färber et al.EMNLP 2025 · 1 citation
- Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNsWenxin Tai, Ting Zhong, Goce Trajcevski, Fan ZhouICLR 2026
- Framework for Evaluating Faithfulness of Local ExplanationsSanjoy Dasgupta, Nave Frost, Michal MoshkovitzICML 2022 · 87 citations
- A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text ExplanationsLingjun Zhao, Hal Daumé IIIEMNLP 2025 · 3 citations
