Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel Simulation
Zechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay, Pulkit Agrawal
Abstract
Reinforcement learning is time-consuming for complex tasks due to the need for large amounts of training data. Recent advances in GPU-based simulation, such as Isaac Gym, have sped up data collection thousands of times on a commodity GPU. Most prior works used on-policy methods like PPO due to their simplicity and ease of scaling. Off-policy methods are more data efficient but challenging to scale, resulting in a longer wall-clock training time. This paper presents a Parallel -Learning (PQL) scheme that outperforms PPO in wall-clock time while maintaining superior sample efficiency of off-policy learning. PQL achieves this by parallelizing data collection, policy learning, and value learning. Different from prior works on distributed off-policy learning, such as Apex, our scheme is designed specifically for massively parallel GPU-based simulation and optimized to work on a single workstation. In experiments, we demonstrate that -learning can be scaled to tens of thousands of parallel environments and investigate important factors affecting learning speed. The code is available at https://github.com/Improbable-AI/pql.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Horizon Reduction Makes RL ScalableSeohong Park, Kevin Frans, Deepinder Mann, Benjamin Eysenbach et al.NeurIPS 2025 · 60 citations
- GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningShutong Ding, Ke Hu, Shan Zhong, Haoyang Luo et al.NeurIPS 2025 · 22 citations
- SAPG: Split and Aggregate Policy GradientsJayesh Singla, Ananye Agarwal, Deepak PathakICML 2024 · 19 citations
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto et al.ICLR 2026 · 12 citations
- Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Hao Wang, Devesh Tiwari, Jian Li et al.SC 2024 · 10 citations
Builds on5
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
- Overcoming The Spectral Bias of Neural Value ApproximationGe Yang, Anurag Ajay, Pulkit AgrawalICLR 2022 · 30 citations
Related papers
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 6 citations
- Differentiable Model Predictive Control on the GPUEmre Adabag, Marcus Greiff, John Subosits, Thomas Jonathan LewICLR 2026 · 13 citations
- FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous PlatformsSamuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer et al.HPDC 2026 · 1 citation
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
- Enhancing Diversity In Parallel Agents: A Maximum State Entropy Exploration StoryVincenzo De Paola, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICML 2025
