Probabilistic Soundness Guarantees in LLM Reasoning Chains
Weiqiu You, Anton Xue, Shreya Havaldar, Delip Rao, Helen Jin, Chris Callison-Burch, Eric Wong
Abstract
In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning. To better detect such errors, we introduce Autoregressive Reasoning Entailment Stability (ARES), a probabilistic framework that evaluates each reasoning step based solely on previously-verified premises. This inductive method yields a nuanced score for each step and provides certified statistical guarantees of its soundness, rather than a brittle binary label. ARES achieves state-of-the-art performance across four benchmarks (72.1% Macro-F1, +8.2 points) and demonstrates superior robustness on very long synthetic reasoning chains, where it excels at detecting propagated errors (90.3% F1, +27.6 points). 1 Correct Reasoning Chain Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 2/5. Claim 4: Therefore, 5x = 6x-14. Claim 5: Finally, we get x = 14. (Correct) Unsound Steps Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 3/5.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8059617e-58a1-4261-985d-ab1e05294510Cited by top-tier papers2
- Automating Database-Native Function Code Synthesis with LLMsWei Zhou, Xuanhe Zhou, Qikang He, Guoliang Li et al.SIGMOD 2026 · 6 citations
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-OzFrederic Gmeiner, Kenneth Holstein, Nikolas MartelaroCHI 2026 · 1 citation
Builds on22
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
Related papers
- Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMsSagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma et al.ICML 2025
- RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning ModelsZihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu et al.ACL 2026
- AutoPSV: Automated Process-Supervised VerifierJianqiao Lu, Zhiyang Dou, Hongru Wang, Zeyu Cao et al.NeurIPS 2024 · 35 citations
- Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal VerificationChengwu Liu, Ye Yuan, Yichun Yin, Yan Xu et al.ACL 2025
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good LensZiyang Ma, Qingyue Yuan, Zhenglin Wang, Deyu ZhouEMNLP 2025 · 7 citations
