Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Hongcheng Wang, Yinuo Huang, Sukai Wang, Guanghui Ren, Hao Dong
Abstract
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce variance, it lacks a theoretical explanation of why it works and whether it is important or potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple continuations are sampled for each thought. Using the multivariate delta method, we reveal a sampling-dimension asymmetry. Increasing sampled thoughts () leaves a strictly positive estimation-variance floor, whereas increasing continuations per thought () drives the leading-order estimation variance to zero at rate . This implies that, within the fixed-temperature GRPO-style estimator without value models studied here, accurate thought-level advantage estimation cannot be achieved by scaling thought sampling alone, making continuation-level branching a principled and potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for its effectiveness and potential necessity, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across vision domains and under different model architectures and sizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 793fc681-5ec7-4e31-b539-18cc0c5b8150Cited by top-tier papers2
- CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image EditingYan Li, Lin Liu, Xiaopeng Zhang, Wei Xue et al.CVPR 2026 · 2 citations
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward DecompositionYanyu Chen, Jiyue Jiang, Dianzhi Yu, Zheng Wu et al.KDD 2026
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen et al.ICLR 2026 · 104 citations
Related papers
- CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning ModelsZhihang Lin, Mingbao Lin, Yuan Xie, Rongrong JiNeurIPS 2025 · 116 citations
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion ModelsYuming Li, Yikai Wang, Yuying Zhu, Zhongyu Zhao et al.ICLR 2026 · 67 citations
- GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language ModelsJixiao Zhang, Chunsheng ZuoEMNLP 2025 · 1 citation
- WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient ReasoningGagan Mundada, Zihan Huang, Rohan Surana, Sheldon Yu et al.ICML 2026
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihoodXingyu Lin, Yilin Wen, Du Su, En Wang et al.ACL 2026
