Probabilistic Soundness Guarantees in LLM Reasoning Chains
Weiqiu You, Anton Xue, Shreya Havaldar, Delip Rao, Helen Jin, Chris Callison-Burch, Eric Wong
摘要
In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning. To better detect such errors, we introduce Autoregressive Reasoning Entailment Stability (ARES), a probabilistic framework that evaluates each reasoning step based solely on previously-verified premises. This inductive method yields a nuanced score for each step and provides certified statistical guarantees of its soundness, rather than a brittle binary label. ARES achieves state-of-the-art performance across four benchmarks (72.1% Macro-F1, +8.2 points) and demonstrates superior robustness on very long synthetic reasoning chains, where it excels at detecting propagated errors (90.3% F1, +27.6 points). 1 Correct Reasoning Chain Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 2/5. Claim 4: Therefore, 5x = 6x-14. Claim 5: Finally, we get x = 14. (Correct) Unsound Steps Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 3/5.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Automating Database-Native Function Code Synthesis with LLMsWei Zhou, Xuanhe Zhou, Qikang He, Guoliang Li 等SIGMOD 2026 · 被引用 6 次
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-OzFrederic Gmeiner, Kenneth Holstein, Nikolas MartelaroCHI 2026 · 被引用 1 次
它引用的顶会 Paper22
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
相关 Paper
- Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMsSagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma 等ICML 2025
- RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning ModelsZihang Liu, Zhouhua Fang, Hui Liu, Zhiwei Liu 等ACL 2026
- AutoPSV: Automated Process-Supervised VerifierJianqiao Lu, Zhiyang Dou, Hongru Wang, Zeyu Cao 等NeurIPS 2024 · 被引用 35 次
- Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal VerificationChengwu Liu, Ye Yuan, Yichun Yin, Yan Xu 等ACL 2025
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good LensZiyang Ma, Qingyue Yuan, Zhenglin Wang, Deyu ZhouEMNLP 2025 · 被引用 7 次
