S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
Muzhi Dai, Chenxu Yang, Qingyi Si
Abstract
As Test-Time Scaling emerges as an active research focus in the large language model community, advanced post-training methods increasingly emphasize extending chain-of-thought (CoT) generation length, thereby enhancing reasoning capabilities to approach Deepseek R1-like reasoning models. However, recent studies reveal that reasoning models (even Qwen3) consistently exhibit excessive thought redundancy in CoT generation. This overthinking issue arises from the inherent limitations of conventional outcome-reward reinforcement learning, which systematically overlooks the regulation of intermediate reasoning processes. This paper introduces Serial-Group Decaying-Reward Policy Optimization (S-GRPO), a novel reinforcement learning paradigm that enables models to implicitly evaluate the sufficiency of intermediate reasoning steps, thereby facilitating early exit in CoT generation. Unlike GRPO, which samples multiple possible reasoning paths in parallel (parallel group), S-GRPO only samples one reasoning path and serially selects multiple temporal positions from the path to exit thinking and directly generate answers (serial group). For correct answers within a serial group, rewards gradually decrease based on the exit positions along the reasoning path from front to back. This design encourages the model to produce more accurate and concise thoughts, while also incentivizing early thinking termination when appropriate. Empirical evaluations demonstrate that S-GRPO is compatible with state-of-the-art reasoning models, including Qwen3 and Deepseek-distill. Across diverse benchmarks such as GSM8K, AIME 2024, AMC 2023, MATH-500, and GPQA Diamond, S-GRPO achieves a substantial reduction in sequence length (35.4% - 61.1%) while simultaneously improving accuracy (absolute 0.72% - 6.08%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b7338f1-ec81-48de-9998-1ce5d8c80cf0Cited by top-tier papers31
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen et al.NeurIPS 2025 · 177 citations
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level ComputationSangmin Bae, Yujin Kim, Reza Bayat, Sungnyun Kim et al.NeurIPS 2025 · 143 citations
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen et al.ICLR 2026 · 104 citations
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question ReformulationYanqi Dai, Yuxiang Ji, Xiao Zhang, Yong Wang et al.ICLR 2026 · 26 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Self-Evaluation Guided Beam Search for ReasoningYuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao et al.NeurIPS 2023 · 316 citations
Related papers
- Step-GRPO: Internalizing Dynamic Early Exit for Efficient ReasoningBenteng Chen, Weida Wang, Shufei Zhang, Mingbao Lin et al.ACL 2026
- Optimizing Length Compression in Large Reasoning ModelsZhengxiang Cheng, Dongping Chen, Mingyang Fu, Tianyi ZhouACL 2026 · 32 citations
- DRPO: Efficient Reasoning via Decoupled Reward Policy OptimizationGang Li, Yan Chen, Ming Lin, Tianbao YangICLR 2026 · 19 citations
- SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model ReasoningChenzhi Hu, Qinzhe Hu, Yuhang Xu, Junyi Chen et al.ICML 2026 · 2 citations
- Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step EntropyZeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen et al.ICLR 2026 · 35 citations
