Lune

EMNLP2025Top-tier venue

Probabilistic Soundness Guarantees in LLM Reasoning Chains

Weiqiu You, Anton Xue, Shreya Havaldar, Delip Rao, Helen Jin, Chris Callison-Burch, Eric Wong

2025Year
9Citations
2Top-tier citations

Abstract

In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detection methods often fail to detect propagated errors because earlier errors can corrupt judgments of downstream reasoning. To better detect such errors, we introduce Autoregressive Reasoning Entailment Stability (ARES), a probabilistic framework that evaluates each reasoning step based solely on previously-verified premises. This inductive method yields a nuanced score for each step and provides certified statistical guarantees of its soundness, rather than a brittle binary label. ARES achieves state-of-the-art performance across four benchmarks (72.1% Macro-F1, +8.2 points) and demonstrates superior robustness on very long synthetic reasoning chains, where it excels at detecting propagated errors (90.3% F1, +27.6 points). 1 Correct Reasoning Chain Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 2/5. Claim 4: Therefore, 5x = 6x-14. Claim 5: Finally, we get x = 14. (Correct) Unsound Steps Claim 1: Let the numerator be x. Claim 2: The denominator is 3x-7. Claim 3: We know that x/(3x-7) = 3/5.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8059617e-58a1-4261-985d-ab1e05294510

Cited by top-tier papers2

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines