S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
Ruotian Ma, Peisong Wang, Cheng Liu, Xingyan Liu, Jiaqi Chen, Bang Zhang, Xin Zhou, Nan Du, Jia Li
摘要
Recent studies have demonstrated the effectiveness of LLM test-time scaling. However, existing approaches to incentivize LLMs' deep thinking abilities generally require large-scale data or significant training efforts. Meanwhile, it remains unclear how to improve the thinking abilities of less powerful base models. In this work, we introduce SR, an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference. Specifically, we first initialize LLMs with iterative self-verification and self-correction behaviors through supervised fine-tuning on carefully curated data. The self-verification and self-correction skills are then further strengthened by both outcome-level and process-level reinforcement learning, with minimized resource requirements, enabling the model to adaptively refine its reasoning process during inference. Our results demonstrate that, with only 3.1k self-verifying and self-correcting behavior initialization samples, Qwen2.5-math-7B achieves an accuracy improvement from 51.0% to 81.6%, outperforming models trained on an equivalent amount of long-CoT distilled data. Extensive experiments and analysis based on three base models across both in-domain and out-of-domain benchmarks validate the effectiveness of SR. Our code and data are available at https://github.com/NineAbyss/S2R.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang 等NeurIPS 2025 · 被引用 85 次
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian 等NeurIPS 2025 · 被引用 69 次
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable RewardsXiaoyuan Liu, Tian Liang, Zhiwei He, Jiahao Xu 等NeurIPS 2025 · 被引用 46 次
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM ReasoningJiaqi Chen, Bang Zhang, Ruotian Ma, Peisong Wang 等NeurIPS 2025 · 被引用 43 次
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- Re2: Unlocking LLM Reasoning via Reinforcement Learning with Re-solvingPinzheng Wang, Shuli Xu, Juntao Li, Yu Luo 等ICLR 2026 · 被引用 7 次
- ReVISE: Learning to Refine at Test-Time via Intrinsic Self-VerificationHyunseok Lee, Seunghyuk Oh, Jaehyung Kim, Jinwoo Shin 等ICML 2025
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu 等NeurIPS 2025 · 被引用 30 次
