High-Throughput Synchronous Deep RL
Iou-Jen Liu, Raymond A. Yeh, Alexander G. Schwing
Abstract
Deep reinforcement learning (RL) is computationally demanding and requires processing of many data points. Synchronous methods enjoy training stability while having lower data throughput. In contrast, asynchronous methods achieve high throughput but suffer from stability issues and lower sample efficiency due to stale policies.' To combine the advantages of both methods we propose High-Throughput Synchronous Deep Reinforcement Learning (HTS-RL). In HTS-RL, we perform learning and rollouts concurrently, devise a system design which avoids stale policies' and ensure that actors interact with environment replicas in an asynchronous manner while maintaining full determinism. We evaluate our approach on Atari games and the Google Research Football environment. Compared to synchronous baselines, HTS-RL is 2-6 faster. Compared to state-of-the-art asynchronous methods, HTS-RL has competitive throughput and consistently achieves higher average episode rewards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Cooperative Exploration for Multi-Agent Deep Reinforcement LearningIou-Jen Liu, Unnat Jain, Raymond A. Yeh, Alexander G. SchwingICML 2021 · 133 citations
- The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal NavigationXiaoming Zhao, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2021 · 50 citations
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 26 citations
- GridToPix: Training Embodied Agents with Minimal SupervisionUnnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi et al.ICCV 2021 · 25 citations
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied RearrangementErik Wijmans, Irfan Essa, Dhruv BatraNeurIPS 2022 · 24 citations
Builds on3
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac et al.AAAI 2020 · 496 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
Related papers
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu et al.NeurIPS 2025 · 273 citations
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 6 citations
- An Efficient Asynchronous Method for Integrating Evolutionary and Gradient-based Policy SearchKyunghyun Lee, Byeong-Uk Lee, Ukcheol Shin, In So KweonNeurIPS 2020 · 24 citations
- Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Hao Wang, Devesh Tiwari, Jian Li et al.SC 2024 · 10 citations
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang et al.EuroSys 2026 · 2 citations
