SAPG: Split and Aggregate Policy Gradients
Jayesh Singla, Ananye Agarwal, Deepak Pathak
Abstract
Despite extreme sample inefficiency, on-policy reinforcement learning, aka policy gradients, has become a fundamental tool in decision-making problems. With the recent advances in GPU-driven simulation, the ability to collect large amounts of data for RL training has scaled exponentially. However, we show that current RL methods, e.g. PPO, fail to ingest the benefit of parallelized environments beyond a certain point and their performance saturates. To address this, we propose a new on-policy RL algorithm that can effectively leverage large-scale environments by splitting them into chunks and fusing them back together via importance sampling. Our algorithm, termed SAPG, shows significantly higher performance across a variety of challenging environments where vanilla PPO and other strong baselines fail to achieve high performance. Website at https://sapg-rl.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e54346ae-be6d-43cd-ac2a-23fe11e2aa3fCited by top-tier papers11
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach et al.NeurIPS 2025 · 60 citations
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino et al.ICLR 2026 · 15 citations
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto et al.ICLR 2026 · 12 citations
- Compute-Optimal Scaling for Value-Based Deep RLPreston Fu, Oleh Rybkin, Zhiyuan Zhou, Michal Nauman et al.NeurIPS 2025 · 7 citations
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 6 citations
Builds on2
Related papers
- Ranking Policy GradientKaixiang Lin, Jiayu ZhouICLR 2020 · 8 citations
- Generalized Proximal Policy Optimization with Sample ReuseJames Queeney, Yannis Paschalidis, Christos G. CassandrasNeurIPS 2021 · 80 citations
- Stabilizing Reinforcement Learning in Differentiable Multiphysics SimulationEliot Xing, Vernon Luk, Jean OhICLR 2025
- Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement LearningNaoki Shitanda, Motoki Omura, Tatsuya Harada, Takayuki OsaICLR 2026
- Enhancing Diversity In Parallel Agents: A Maximum State Entropy Exploration StoryVincenzo De Paola, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICML 2025
