Large Language Models Can Self-Correct with Key Condition Verification
Zhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan, Chao Shen, Meng Jiang
Abstract
Intrinsic self-correct was a method that instructed large language models (LLMs) to verify and correct their responses without external feedback. Unfortunately, the study concluded that the LLMs could not self-correct reasoning yet. We find that a simple yet effective prompting method enhances LLM performance in identifying and correcting inaccurate answers without external feedback. That is to mask a key condition in the question, add the current response to construct a verification question, and predict the condition to verify the response. The condition can be an entity in an open-domain question or a numerical value in an arithmetic question, which requires minimal effort (via prompting) to identify. We propose an iterative verify-then-correct framework to progressively identify and correct (probably) false responses, named PROCO. We conduct experiments on three reasoning tasks. On average, PROCO, with GPT-3.5-Turbo-1106 as the backend LLM, yields +6.8 exact match on four open-domain question answering datasets, +14.1 accuracy on three arithmetic reasoning datasets, and +9.6 accuracy on a commonsense reasoning dataset, compared to Self-Correct. Our implementation is made publicly available at https://wzy6642.github.io/proco. github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e726cc0a-d3b6-4e50-8589-b6f4d5e44d83Cited by top-tier papers12
- Understanding the Dark Side of LLMs' Intrinsic Self-CorrectionQingjie Zhang, Di Wang, Haoting Qian, Yiming Li et al.ACL 2025 · 36 citations
- Tina: Tiny Reasoning Models via LoRAShangshang Wang, Julian Asilis, Ömer Faruk Akgül, Enes Burak Bilgin et al.ICLR 2026 · 30 citations
- Controlling Thinking Speed in Reasoning ModelsZhengkai Lin, Zhihang Fu, Ze Chen, Chao Chen et al.NeurIPS 2025 · 20 citations
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External FeedbackDongwei Jiang, Bowei Zhang, Andrew Wang, Nicholas Andrews et al.NeurIPS 2025 · 11 citations
- SciCoQA: Quality Assurance for Scientific Paper-Code AlignmentTim Baumgärtner, Iryna GurevychACL 2026 · 5 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
Related papers
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Enhancing Mathematical Reasoning in LLMs by Stepwise CorrectionZhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan et al.ACL 2025
- Small Language Model Can Self-CorrectHaixia Han, Jiaqing Liang, Jie Shi, Qianyu He et al.AAAI 2024 · 31 citations
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningNing Miao, Yee Whye Teh, Tom RainforthICLR 2024 · 195 citations
- Get an A in Math: Progressive Rectification PromptingZhenyu Wu, Meng Jiang, Chao ShenAAAI 2024 · 15 citations
