Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
Shangyu Xing, Siyuan Wang, Chenyuan Yang, Xin-Yu Dai, Xiang Ren
摘要
Reinforcement Learning with Verifiable Rewards (RLVR), particularly with algorithms like Group Relative Policy Optimization (GRPO), has proven highly effective in enhancing the reasoning capabilities of large language models. However, a critical bottleneck in current pipelines lies in the limited diversity of sampled trajectories during group rollouts. Homogeneous trajectories and their associated rewards would diminish the return signals for policy updates, thereby hindering effective policy learning. This lack of diversity stems primarily from token-level stochastic sampling, where local variations are likely to collapse into near-identical reasoning paths. To address this limitation, we propose Lookahead Tree-Based Rollouts (LATR), a novel rollout strategy designed to explicitly promotes trajectory-level diversity by enforcing branching into different candidate tokens likely to yield distinct continuations. Specifically, LATR iteratively operates in three stages: (1) branching at high-uncertainty generation steps, (2) performing lookahead simulation for each new branch, and (3) pruning branches that exhibits prolonged similarity during simulation. Compared with stochastic Sampling, LATR accelerates policy learning by 131% on average and improves final pass@1 performance by 4.2% on both GRPO and Dynamic sAmpling Policy Optimization (DAPO) algorithms across different reasoning tasks. Our code and data are publicly available at https://github.com/starreeze/latr.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic LearningJinyang Wu, Shuo Yang, Yuhao Shen, Shuai Zhang 等ACL 2026 · 被引用 9 次
- Segment-Level Attribution for Selective Learning of Long Reasoning TracesSiyuan Wang, Yanchen Liu, Xiang RenICLR 2026 · 被引用 2 次
它引用的顶会 Paper10
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu 等NeurIPS 2025 · 被引用 181 次
相关 Paper
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive ExplorationZhicheng Yang, Zhijiang Guo, Yinya Huang, Yongxin Wang 等ICML 2026 · 被引用 38 次
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsHaoran He, Yuxiao Ye, Qingpeng Cai, Chen Hu 等ICLR 2026 · 被引用 9 次
- I²B-LPO: Latent Policy Optimization via Iterative Information BottleneckHuilin Deng, Hongchen Luo, Yue Zhu, Long Li 等ACL 2026
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVRHaobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan 等ACL 2026 · 被引用 6 次
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu 等ICLR 2026 · 被引用 51 次
