Large Batch Simulation for Deep Reinforcement Learning
Brennan Shacklett, Erik Wijmans, Aleksei Petrenko, Manolis Savva, Dhruv Batra, Vladlen Koltun, Kayvon Fatahalian
Abstract
We accelerate deep reinforcement learning-based training in visually complex 3D environments by two orders of magnitude over prior work, realizing end-to-end training speeds of over 19,000 frames of experience per second on a single GPU and up to 72,000 frames per second on a single eight-GPU machine. The key idea of our approach is to design a 3D renderer and embodied navigation simulator around the principle of "batch simulation": accepting and executing large batches of requests simultaneously. Beyond exposing large amounts of work at once, batch simulation allows implementations to amortize in-memory storage of scene assets, rendering work, data loading, and synchronization costs across many simulation requests, dramatically improving the number of simulated agents per GPU and overall simulation throughput. To balance DNN inference and training costs with faster simulation, we also build a computationally efficient policy DNN that maintains high task performance, and modify training algorithms to maintain sample efficiency when training with large mini-batches. By combining batch simulation and DNN performance optimizations, we demonstrate that PointGoal navigation agents can be trained in complex 3D environments on a single GPU in 1.5 days to 97% of the accuracy of agents trained on a prior state-of-the-art system using a 64-GPU cluster over three days. We provide open-source reference implementations of our batch 3D renderer and simulator to facilitate incorporation of these ideas into RL systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbb793cb-481f-4cc4-97e2-b7400d57515eCited by top-tier papers6
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Megaverse: Simulating Embodied Agents at One Million Experiences per SecondAleksei Petrenko, Erik Wijmans, Brennan Shacklett, Vladlen KoltunICML 2021 · 26 citations
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied RearrangementErik Wijmans, Irfan Essa, Dhruv BatraNeurIPS 2022 · 24 citations
- An Extensible, Data-Oriented Architecture for High-Performance, Many-World SimulationBrennan Shacklett, Luc Guy Rosenzweig, Zhiqiang Xie, Bidipta Sarkar et al.SIGGRAPH 2023 · 13 citations
- Galactic: Scaling End-to-End Reinforcement Learning for Rearrangement at 100k Steps-Per-SecondVincent-Pierre Berges, Andrew Szot, Devendra Singh Chaplot, Aaron Gokaslan et al.CVPR 2023
Builds on7
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme et al.ICML 2020 · 131 citations
- Accelerating Reinforcement Learning through GPU Atari EmulationSteven Dalton, Iuri FrosioNeurIPS 2020 · 50 citations
Related papers
- GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPSSaman Kazemkhani, Aarav Pandya, Daphne Cornelisse, Brennan Shacklett et al.ICLR 2025
- Accelerating Visual-Policy Learning through Parallel Differentiable SimulationHaoxiang You, Yilang Liu, Ian AbrahamNeurIPS 2025 · 7 citations
- Influence-Augmented Local Simulators: a Scalable Solution for Fast Deep RL in Large Networked SystemsMiguel Suau, Jinke He, Matthijs T. J. Spaan, Frans A. OliehoekICML 2022 · 5 citations
- TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU SystemsYing Li, Yuhui Bao, Gongyu Wang, Xinxin Mei et al.ISCA 2025 · 2 citations
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen et al.USENIX ATC 2023 · 9 citations
