Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
Mingkang Zhu, Xi Chen, Bei Yu, Hengshuang Zhao, Jiaya Jia
摘要
Large language model (LLM) agents increasingly rely on external tools such as search engines to solve complex, multi-step problems, and reinforcement learning (RL) has become a key paradigm for training them. However, the trajectories of search agents are structurally heterogeneous, where variations in the number, placement, and outcomes of search calls lead to fundamentally different answer directions and reward distributions. Standard policy gradient methods, which use a single global baseline, suffer from what we identify and formalize as crossstratum bias-an "apples-to-oranges" comparison of heterogeneous trajectories. This cross-stratum bias distorts credit assignment and hinders exploration of complex, multi-step search strategies. To address this, we propose Stratified GRPO, whose central component, Stratified Advantage Normalization (SAN), partitions trajectories into homogeneous strata based on their structural properties and computes advantages locally within each stratum. This ensures that trajectories are evaluated only against their true peers. Our analysis proves that SAN eliminates cross-stratum bias, yields conditionally unbiased unit-variance estimates inside each stratum, and retains the global unbiasedness and unit-variance properties enjoyed by standard normalization, resulting in a more pure and scale-stable learning signal. To improve practical stability under finite-sample regimes, we further linearly blend SAN with the global estimator. Extensive experiments on diverse single-hop and multi-hop question-answering benchmarks demonstrate that Stratified GRPO consistently and substantially outperforms GRPO by up to 11.3 points, achieving higher training rewards, greater training stability, and more effective search policies. These results establish stratification as a principled remedy for structural heterogeneity in RL for LLM search agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PDCR: Perception-Decomposed Confidence Reward for Vision-Language ReasoningHee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom 等CVPR 2026 · 被引用 3 次
- VisionLeaf: Entropy-Guided Leaf-First Reasoning for Efficient and Accurate Think-with-ImageHaokun GUI, Senqiao Yang, Mingkang Zhu, Meng Chu 等CVPR 2026
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- Tree Search for LLM Agent Reinforcement LearningYuxiang Ji, Ziyu Ma, Yong Wang, Guanhua Chen 等ICLR 2026 · 被引用 71 次
- Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationJunzhe Wang, Zhiheng Xi, Yajie Yang, Hao Luo 等ACL 2026 · 被引用 4 次
- STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent TrainingQiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang 等ACL 2026
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang 等NeurIPS 2025 · 被引用 3 次
