Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
Chaorui Yao, Yanxi Chen, Yuchang Sun, Yushuo Chen, Wenhao Zhang, Xuchen Pan, Yaliang Li, Bolin Ding
摘要
Off-policy reinforcement learning (RL) for large language models (LLMs) is attracting growing interest, driven by practical constraints in real-world applications, the complexity of LLM-RL infrastructure, and the need for further innovations of RL methodologies. While classic REINFORCE and its modern variants like Group Relative Policy Optimization (GRPO) are typically regarded as on-policy algorithms with limited tolerance of off-policyness, we present in this work a first-principles derivation for group-relative REINFORCE -a REINFORCE variant that uses the within-group mean reward as the baseline for advantage calculation -without assuming a specific training data distribution, showing that it admits a native off-policy interpretation. This perspective yields two general principles for adapting REINFORCE to truly off-policy settings: regularizing policy updates, and actively shaping the data distribution. Our analysis demystifies some myths about the roles of importance sampling and clipping in GRPO, unifies and reinterprets two recent algorithms -Online Policy Mirror Descent and Asymmetric REINFORCE -as regularized forms of the REINFORCE loss, and offers theoretical justification for seemingly heuristic dataweighting strategies. Our findings lead to actionable insights that are validated with extensive empirical studies, and open up new opportunities for principled algorithm design in off-policy RL for LLMs. Source code for this work is available at https://github.com/agentscope-ai/Trinity-RFT/ tree/main/examples/rec_gsm8k . * Equal contribution. Part of the work was done while Chaorui Yao was an intern at Alibaba Group.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy DataFahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov 等ICML 2024 · 被引用 189 次
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
相关 Paper
- Reinforcement Learning with Verifiable Rewards: GRPO's Loss, Dynamics, and Success AmplificationYoussef MrouehICML 2026 · 被引用 118 次
- GVPO: Group Variance Policy Optimization for Large Language Model Post-TrainingKaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang 等NeurIPS 2025 · 被引用 35 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashionYannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi 等EMNLP 2024
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewardsCharles Arnal, Gaëtan Narozniak, Vivien Cabannes, Yunhao Tang 等NeurIPS 2025 · 被引用 30 次
