Accelerating Reinforcement Learning through GPU Atari Emulation
Steven Dalton, Iuri Frosio
Abstract
We introduce CuLE (CUDA Learning Environment), a CUDA port of the Atari Learning Environment (ALE) which is used for the development of deep reinforcement algorithms. CuLE overcomes many limitations of existing CPU-based emulators and scales naturally to multiple GPUs. It leverages GPU parallelization to run thousands of games simultaneously and it renders frames directly on the GPU, to avoid the bottleneck arising from the limited CPU-GPU communication bandwidth. CuLE generates up to 155M frames per hour on a single GPU, a finding previously achieved only through a cluster of CPUs. Beyond highlighting the differences between CPU and GPU emulators in the context of reinforcement learning, we show how to leverage the high throughput of CuLE by effective batching of the training data, and show accelerated convergence for A2C+V-trace. CuLE is available at https://github.com/NVlabs/cule .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6305688-dc4f-4d84-9c3e-69b55f698782Cited by top-tier papers14
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAXClément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana et al.ICLR 2024 · 52 citations
- Atari-5: Distilling the Arcade Learning Environment down to Five GamesMatthew Aitchison, Penny Sweetser, Marcus HutterICML 2023 · 40 citations
- Large Batch Simulation for Deep Reinforcement LearningBrennan Shacklett, Erik Wijmans, Aleksei Petrenko, Manolis Savva et al.ICLR 2021 · 29 citations
- An Extensible, Data-Oriented Architecture for High-Performance, Many-World SimulationBrennan Shacklett, Luc Guy Rosenzweig, Zhiqiang Xie, Bidipta Sarkar et al.SIGGRAPH 2023 · 13 citations
Builds on1
Related papers
- Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAXWaris Radji, Thomas Michel, Hector PiteauICLR 2026 · 5 citations
- Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel SimulationZechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay et al.ICML 2023 · 27 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang et al.VLDB 2025 · 1 citation
- High-Throughput Synchronous Deep RLIou-Jen Liu, Raymond A. Yeh, Alexander G. SchwingNeurIPS 2020 · 13 citations
