A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva
Abstract
Prompting language models to provide stepby-step answers (e.g., "Chain-of-Thought") is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance. Recent literature discusses automatic methods to verify reasoning steps to evaluate and improve their correctness. However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction. We introduce REVEAL: Reasoning Verification Evaluation, a new dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question answering settings. REVEAL includes comprehensive labels for the relevance, attribution to evidence passages, and logical correctness of each reasoning step in a language model's answer, across a wide variety of datasets and state-of-the-art language models. Available at reveal-dataset.github.io. Which population is bigger, California or Texas? (1) Texas population is 39.53m. (2) California population is 39.24m. (3) Thus, California has a bigger population than Texas. (4) So the answer is California.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dc3ab7f-1877-4056-80c7-0f02fc5c574aCited by top-tier papers22
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 37 citations
- MiniCheck: Efficient Fact-Checking of LLMs on Grounding DocumentsLiyan Tang, Philippe Laban, Greg DurrettEMNLP 2024 · 26 citations
- Verifying Chain-of-Thought Reasoning via Its Computational GraphZheng Zhao, Yeskendir Koishekenov, Xianjun Yang, Naila Murray et al.ICLR 2026 · 25 citations
- Privacy Reasoning in Ambiguous ContextsRen Yi, Octavian Suciu, Adrià Gascón, Sarah Meiklejohn et al.NeurIPS 2025 · 15 citations
- Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationSanjana Ramprasad, Byron C. WallaceNeurIPS 2025 · 13 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
Related papers
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
- STREET: A Multi-Task Structured Reasoning and Explanation BenchmarkDanilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Henghui Zhu et al.ICLR 2023 · 6 citations
- CHECKWHY: Causal Fact Verification via Argument StructureJiasheng Si, Yibo Zhao, Yingjie Zhu, Haiyang Zhu et al.ACL 2024
- Large Language Models Meet Symbolic Provers for Logical Reasoning EvaluationChengwen Qi, Ren Ma, Bowen Li, He Du et al.ICLR 2025
- MetaLogic: Logical Reasoning Explanations with Fine-Grained StructureYinya Huang, Hongming Zhang, Ruixin Hong, Xiaodan Liang et al.EMNLP 2022 · 3 citations
