ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness
Archiki Prasad, Swarnadeep Saha, Xiang Zhou, Mohit Bansal
摘要
Multi-step reasoning ability is fundamental to many natural language tasks, yet it is unclear what constitutes a good reasoning chain and how to evaluate them. Most existing methods focus solely on whether the reasoning chain leads to the correct conclusion, but this answeroriented view may confound reasoning quality with other spurious shortcuts to predict the answer. To bridge this gap, we evaluate reasoning chains by viewing them as informal proofs that derive the final answer. Specifically, we propose RECEVAL (Reasoning Chain Evaluation), a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, i.e., each step provides new information that is helpful towards deriving the generated answer. We evaluate these properties by developing metrics using natural language inference models and V-Information. On multiple datasets, we show that RECEVAL effectively identifies various error types and yields notable improvements compared to prior methods. We analyze the impact of step boundaries, and previous steps on evaluating correctness and demonstrate that our informativeness metric captures the expected flow of information in high-quality reasoning chains. Finally, we show that scoring reasoning chains based on RECEVAL improves downstream task performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang 等NeurIPS 2023 · 被引用 234 次
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li 等NeurIPS 2024 · 被引用 51 次
- Advancing Process Verification for Large Language Models via Tree-Based Preference LearningMingqian He, Yongliang Shen, Wenqi Zhang, Zeqi Tan 等EMNLP 2024 · 被引用 14 次
- Probabilistic Soundness Guarantees in LLM Reasoning ChainsWeiqiu You, Anton Xue, Shreya Havaldar, Delip Rao 等EMNLP 2025 · 被引用 9 次
- A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning ChainsAlon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig 等ACL 2024 · 被引用 8 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
相关 Paper
- REV: Information-Theoretic Evaluation of Free-Text RationalesHanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji 等ACL 2023 · 被引用 14 次
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step ReasoningOlga Golovneva, Moya Chen, Spencer Poff, Martin Corredor 等ICLR 2023 · 被引用 28 次
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question AnsweringFrancesco Maria Molfese, Luca Moroni, Ciro Porcaro, Simone Conia 等ACL 2026 · 被引用 1 次
- ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain ExtractionPengze Li, Jiaqi Liu, Junchi Yu, Lihao Liu 等AAAI 2026 · 被引用 1 次
