Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
Sagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma, Dilek Hakkani-Tur
Abstract
Chain-of-Thought (CoT) prompting enhances mathematical reasoning in large language models (LLMs) by enabling detailed step-by-step solutions. However, due to the verbosity of LLMs, the resulting reasoning chains can be long, making it harder to verify the reasoning steps and trace issues resulting from dependencies between the steps that may be farther away in the sequence of steps. Importantly, mathematical reasoning allows each step to be derived from a small set of premises, which are a subset of the preceding steps in the reasoning chain. In this paper, we present a framework that identifies the premises for each step, to improve the evaluation of reasoning. We restructure conventional linear reasoning chains into Premise Augmented Reasoning Chains (PARC) by introducing premise links, resulting in a directed acyclic graph where the nodes are the steps and the edges are the premise links. Through experiments with a PARC-based dataset that we built, namely PERL (Premises and ERrors identification in LLMs), we demonstrate that LLMs can reliably identify premises within complex reasoning chains. In particular, even open-source LLMs achieve 90% recall in premise identification. We also show that PARC helps to identify errors in reasoning chains more reliably. The accuracy of error identification improves by 6% to 16% absolute when step-bystep verification is carried out in PARC under the premises. Our findings highlight the utility of premise-centric representations in addressing complex problem-solving tasks and open new avenues for improving the reliability of LLM-based reasoning evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ae3f840-3bc2-467d-92af-59e7797c09a2Cited by top-tier papers7
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to DeliberationZhiwei Zhang, Xiaomin Li, Yudi Lin, Hui Liu et al.ICLR 2026 · 13 citations
- Probabilistic Soundness Guarantees in LLM Reasoning ChainsWeiqiu You, Anton Xue, Shreya Havaldar, Delip Rao et al.EMNLP 2025 · 9 citations
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic ConfidenceAmirhosein Ghasemabadi, Keith G. Mills, Baochun Li, Di NiuACL 2026 · 9 citations
- Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence AttributionXiaoou Liu, Tiejin Chen, Dengjia Zhang, Yaqing Wang et al.ICML 2026 · 1 citation
- Hallucination Detection from Structural Reasoning ModelJianbo Sun, Pengkun YangICML 2026
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
Related papers
- Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsYifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng et al.EMNLP 2024 · 3 citations
- Get an A in Math: Progressive Rectification PromptingZhenyu Wu, Meng Jiang, Chao ShenAAAI 2024 · 15 citations
- Causal Sufficiency and Necessity Improves Chain-of-Thought ReasoningXiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li et al.NeurIPS 2025 · 19 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
- Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What MattersBoshi Wang, Sewon Min, Xiang Deng, Jiaming Shen et al.ACL 2023 · 100 citations
