Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving
Xin Quan, Marco Valentino, Louise A. Dennis, André Freitas
Abstract
Natural language explanations represent a proxy for evaluating explanation-based and multi-step Natural Language Inference (NLI) models. However, assessing the validity of explanations for NLI is challenging as it typically involves the crowd-sourcing of apposite datasets, a process that is time-consuming and prone to logical errors. To address existing limitations, this paper investigates the verification and refinement of natural language explanations through the integration of Large Language Models (LLMs) and Theorem Provers (TPs). Specifically, we present a neuro-symbolic framework, named Explanation-Refiner, that integrates TPs with LLMs to generate and formalise explanatory sentences and suggest potential inference strategies for NLI. In turn, the TP is employed to provide formal guarantees on the logical validity of the explanations and to generate feedback for subsequent improvements. We demonstrate how Explanation-Refiner can be jointly used to evaluate explanatory reasoning, autoformalisation, and error correction mechanisms of state-of-the-art LLMs as well as to automatically enhance the quality of explanations of variable complexity in different domains. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52449fb5-17f8-401f-98a2-1e6ce6ed745fCited by top-tier papers16
- VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency ChecksYu Feng, Nathaniel Weir, Kaj Bostrom, Sam Bayless et al.ICLR 2026 · 16 citations
- Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation SteeringMarco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao et al.AAAI 2026 · 15 citations
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task PlanningSanghyun Ahn, Wonje Choi, Junyong Lee, Jinwoo Park et al.NeurIPS 2025 · 14 citations
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning TasksDebargha Ganguly, Vikash Singh, Sreehari Sankar, Biyao Zhang et al.NeurIPS 2025 · 11 citations
- LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical SupervisionJundong Xu, Hao (Scofield) Fei, Huichi Zhou, Xin Quan et al.ICLR 2026 · 9 citations
Builds on14
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- Baldur: Whole-Proof Generation and Repair with Large Language ModelsEmily First, Markus N. Rabe, Talia Ringer, Yuriy BrunFSE 2023 · 89 citations
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversTheo Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang et al.EMNLP 2023 · 37 citations
Related papers
- Faithful and Robust LLM-Driven Theorem Proving for NLI ExplanationsXin Quan, Marco Valentino, Louise A. Dennis, André FreitasACL 2025 · 8 citations
- Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem ProvingXinyi Zheng, Ningke Li, Xiaokun Luan, Wang Kailong et al.ICSE 2026
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- NeSTR: A Neuro-Symbolic Abductive Framework for Temporal Reasoning in Large Language ModelsFeng Liang, Weixin Zeng, Runhao Zhao, Xiang ZhaoAAAI 2026
- NaturalProver: Grounded Mathematical Proof Generation with Language ModelsSean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi et al.NeurIPS 2022 · 108 citations
