Single-stream Policy Optimization
Zhongwen Xu, Zihan Ding
Abstract
We revisit policy-gradient optimization for Large Language Models (LLMs) from a single-stream perspective. Prevailing group-based methods like GRPO reduce variance with on-the-fly baselines but suffer from critical flaws: frequent degenerate groups erase learning signals, and synchronization barriers hinder scalability. We introduce Single-stream Policy Optimization (SPO), which eliminates these issues by design. SPO replaces per-group baselines with a persistent, KL-adaptive value tracker and normalizes advantages globally across the batch, providing a stable, low-variance learning signal for every sample. Being group-free, SPO enables higher throughput and scales effectively in long-horizon or tool-integrated settings where generation times vary. Furthermore, the persistent value tracker naturally enables an adaptive curriculum via prioritized sampling. Experiments using Qwen3-8B show that SPO converges more smoothly and attains higher accuracy than GRPO, while eliminating computation wasted on degenerate groups. Ablation studies confirm that SPO's gains stem from its principled approach to baseline estimation and advantage normalization, offering a more robust and efficient path for LLM reasoning. Across five hard math benchmarks with Qwen3-8B, SPO improves the average maj@32 by over GRPO, driven by substantial absolute point gains on challenging datasets, including on BRUMO 25, on AIME 25, on HMMT 25, and achieves consistent relative gain in pass@ across the evaluated values. SPO's success challenges the prevailing trend of adding incidental complexity to RL algorithms, highlighting a path where fundamental principles, not architectural workarounds, drive the next wave of progress in LLM reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fab187ec-eae0-4d2a-9895-03db1c438fabCited by top-tier papers7
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement FinetuningQianli Shen, Daoyuan Chen, Yilun Huang, Zhenqing Ling et al.ICLR 2026 · 15 citations
- Stable and Efficient Single-Rollout RL for Multimodal ReasoningRui Liu, Dian Yu, Lei Ke, Haolin Liu et al.CVPR 2026 · 13 citations
- Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable RewardsGuanning Zeng, Zhaoyi Zhou, Daman Arora, Andrea ZanetteICML 2026 · 13 citations
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsHaoran He, Yuxiao Ye, Qingpeng Cai, Chen Hu et al.ICLR 2026 · 9 citations
- Segment-Aligned Policy Optimization for Multi-Modal ReasoningLei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan et al.ICML 2026 · 2 citations
Builds on12
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian et al.ICLR 2026 · 171 citations
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsHaizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura et al.NeurIPS 2025 · 125 citations
Related papers
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language ModelsYiran Guo, Lijie Xu, Ji Liu, Dan Ye et al.NeurIPS 2025 · 75 citations
- Advantage Collapse in Group Relative Policy Optimization: Diagnosis and MitigationXixiang He, Qiyao Sun, Ao Cheng, Xingming Li et al.ICML 2026
- PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy OptimizationXinxin Zhu, Ying He, Haowen Hou, Ruichong Zhang et al.AAAI 2026
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM ReasoningYifan Zhang, Yifeng Liu, Rina Hughes, Yang Yuan et al.ICLR 2026 · 30 citations
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng et al.ICLR 2026 · 4 citations
