Self-Critique and Refinement for Faithful Natural Language Explanations
Yingming Wang, Pepa Atanasova
Abstract
With the rapid development of Large Language Models (LLMs), Natural Language Explanations (NLEs) have become increasingly important for understanding model predictions. However, these explanations often fail to faithfully represent the model's actual reasoning process. While existing work has demonstrated that LLMs can self-critique and refine their initial outputs for various tasks, this capability remains unexplored for improving explanation faithfulness. To address this gap, we introduce Self-critique and Refinement for Natural Language Explanations (SR-NLE), a framework that enables models to improve the faithfulness of their own explanations -specifically, posthoc NLEs -through an iterative critique and refinement process without external supervision. Our framework leverages different feedback mechanisms to guide the refinement process, including natural language self-feedback and, notably, a novel feedback approach based on feature attribution that highlights important input words. Our experiments across three datasets and four state-of-the-art LLMs demonstrate that SR-NLE significantly reduces unfaithfulness rates, with our best method achieving an average unfaithfulness rate of 36.02%, compared to 54.81% for baseline -an absolute reduction of 18.79%. These findings reveal that the investigated LLMs can indeed refine their explanations to better reflect their actual reasoning process, requiring only appropriate guidance through feedback without additional training or fine-tuning. Our code is available at https://github.com/ymwangv/SR-NLE . Identify the logical relationship between premise and hypothesis. Premise: A man in a red shirt is playing guitar on stage. Hypothesis: A man is performing music. Answer: Entailment Please, provide an explanation for your answer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be4bc8e3-3d8c-442a-b428-3d919579276bCited by top-tier papers1
Ask how each one uses itBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
- Self-AMPLIFY: Improving Small Language Models with Self Post Hoc ExplanationsMilan Bhan, Jean-Noël Vittaut, Nicolas Chesneau, Marie-Jeanne LesotEMNLP 2024 · 1 citation
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun et al.NeurIPS 2023 · 87 citations
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo et al.VLDB 2026 · 2 citations
- S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical LearningWei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao et al.CVPR 2023
