ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness
Archiki Prasad, Swarnadeep Saha, Xiang Zhou, Mohit Bansal
Abstract
Multi-step reasoning ability is fundamental to many natural language tasks, yet it is unclear what constitutes a good reasoning chain and how to evaluate them. Most existing methods focus solely on whether the reasoning chain leads to the correct conclusion, but this answeroriented view may confound reasoning quality with other spurious shortcuts to predict the answer. To bridge this gap, we evaluate reasoning chains by viewing them as informal proofs that derive the final answer. Specifically, we propose RECEVAL (Reasoning Chain Evaluation), a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, i.e., each step provides new information that is helpful towards deriving the generated answer. We evaluate these properties by developing metrics using natural language inference models and V-Information. On multiple datasets, we show that RECEVAL effectively identifies various error types and yields notable improvements compared to prior methods. We analyze the impact of step boundaries, and previous steps on evaluating correctness and demonstrate that our informativeness metric captures the expected flow of information in high-quality reasoning chains. Finally, we show that scoring reasoning chains based on RECEVAL improves downstream task performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bf314cc-059f-46cf-850d-f5035d9b6cc0Cited by top-tier papers20
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang et al.NeurIPS 2023 · 234 citations
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li et al.NeurIPS 2024 · 51 citations
- Advancing Process Verification for Large Language Models via Tree-Based Preference LearningMingqian He, Yongliang Shen, Wenqi Zhang, Zeqi Tan et al.EMNLP 2024 · 14 citations
- Probabilistic Soundness Guarantees in LLM Reasoning ChainsWeiqiu You, Anton Xue, Shreya Havaldar, Delip Rao et al.EMNLP 2025 · 9 citations
- A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning ChainsAlon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig et al.ACL 2024 · 8 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
Related papers
- REV: Information-Theoretic Evaluation of Free-Text RationalesHanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji et al.ACL 2023 · 14 citations
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step ReasoningOlga Golovneva, Moya Chen, Spencer Poff, Martin Corredor et al.ICLR 2023 · 28 citations
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question AnsweringFrancesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia et al.ACL 2026 · 1 citation
- ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain ExtractionPengze Li, Jiaqi Liu, Junchi Yu, Lihao Liu et al.AAAI 2026 · 1 citation
