Generalized Proximal Policy Optimization with Sample Reuse
James Queeney, Yannis Paschalidis, Christos G. Cassandras
Abstract
In real-world decision making tasks, it is critical for data-driven reinforcement learning methods to be both stable and sample efficient. On-policy methods typically generate reliable policy improvement throughout training, while off-policy methods make more efficient use of data through sample reuse. In this work, we combine the theoretically supported stability benefits of on-policy algorithms with the sample efficiency of off-policy algorithms. We develop policy improvement guarantees that are suitable for the off-policy setting, and connect these bounds to the clipping mechanism used in Proximal Policy Optimization. This motivates an off-policy version of the popular algorithm that we call Generalized Proximal Policy Optimization with Sample Reuse. We demonstrate both theoretically and empirically that our algorithm delivers improved performance by effectively balancing the competing goals of stability and sample efficiency. On-policy reinforcement learning methods such as Proximal Policy Optimization (PPO) [19] deliver stable performance throughout training due to their connection to theoretical policy improvement guarantees. These methods are motivated by a lower bound on the expected performance loss at every update, which can be approximated using samples generated by the current policy. The theoretically supported stability of these methods is very attractive, but the need for on-policy data and the highvariance nature of reinforcement learning often requires significant data to be collected between every update, resulting in high sample complexity and slow learning. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f02b049-4fb9-40ac-a726-49a78ca2edddCited by top-tier papers11
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout ReplayYifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen et al.NeurIPS 2025 · 63 citations
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 41 citations
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy TrainingYoussef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure et al.ICLR 2026 · 39 citations
- An Implicit Trust Region Approach to Behavior Regularized Offline Reinforcement LearningZhe Zhang, Xiaoyang TanAAAI 2024 · 10 citations
- Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language ModelJing Liang, Jinyi Liu, Yi Ma, Hongyao Tang et al.ICLR 2026 · 10 citations
Builds on2
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini et al.ICLR 2021 · 52 citations
Related papers
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
- On Stationary Point Convergence of PPO-ClipRuinan Jin, Shuai Li, Baoxiang WangICLR 2024 · 14 citations
- Rethinking the Trust Region in LLM Reinforcement LearningPenghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang et al.ICML 2026 · 22 citations
- Simple Policy OptimizationZhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter et al.ICML 2025
