Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Hongcheng Wang, Yinuo Huang, Sukai Wang, Guanghui Ren, Hao Dong
摘要
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce variance, it lacks a theoretical explanation of why it works and whether it is important or potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple continuations are sampled for each thought. Using the multivariate delta method, we reveal a sampling-dimension asymmetry. Increasing sampled thoughts () leaves a strictly positive estimation-variance floor, whereas increasing continuations per thought () drives the leading-order estimation variance to zero at rate . This implies that, within the fixed-temperature GRPO-style estimator without value models studied here, accurate thought-level advantage estimation cannot be achieved by scaling thought sampling alone, making continuation-level branching a principled and potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for its effectiveness and potential necessity, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across vision domains and under different model architectures and sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image EditingYan Li, Lin Liu, Xiaopeng Zhang, Wei Xue 等CVPR 2026 · 被引用 2 次
- LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward DecompositionYanyu Chen, Jiyue Jiang, Dianzhi Yu, Zheng Wu 等KDD 2026
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen 等ICLR 2026 · 被引用 104 次
相关 Paper
- CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning ModelsZhihang Lin, Mingbao Lin, Yuan Xie, Rongrong JiNeurIPS 2025 · 被引用 116 次
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion ModelsYuming Li, Yikai Wang, Yuying Zhu, Zhongyu Zhao 等ICLR 2026 · 被引用 67 次
- GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language ModelsJixiao Zhang, Chunsheng ZuoEMNLP 2025 · 被引用 1 次
- WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient ReasoningGagan Mundada, Zihan Huang, Rohan Surana, Sheldon Yu 等ICML 2026
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihoodXingyu Lin, Yilin Wen, Du Su, En Wang 等ACL 2026
