Cleanba: A Reproducible and Efficient Distributed Reinforcement Learning Platform
Shengyi Huang, Jiayi Weng, Rujikorn Charakorn, Min Lin, Zhongwen Xu, Santiago Ontañón
Abstract
Distributed Deep Reinforcement Learning (DRL) aims to leverage more computational resources to train autonomous agents with less training time. Despite recent progress in the field, reproducibility issues have not been sufficiently explored. This paper first shows that the typical actor-learner framework can have reproducibility issues even if hyperparameters are controlled. We then introduce Cleanba, a new open-source platform for distributed DRL that proposes a highly reproducible architecture. Cleanba implements highly optimized distributed variants of PPO (Schulman et al., 2017) and IMPALA (Espeholt et al., 2018). Our Atari experiments show that these variants can obtain equivalent or higher scores than strong IMPALA baselines in moolib and torchbeast and PPO baseline in CleanRL. However, Cleanba variants present 1) shorter training time and 2) more reproducible learning curves in different hardware settings. Cleanba's source code is available at https://github.com/vwxyzjn/cleanba * Currently at OpenAI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e7cb00b-bb85-4756-8e30-21fb19d0a5cbCited by top-tier papers2
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-TrainingBrian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain et al.NeurIPS 2025 · 34 citations
- Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language ModelsMichael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini et al.ICLR 2025
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme et al.ICML 2020 · 131 citations
Related papers
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target NetworksMichael Luo, Jiahao Yao, Richard Liaw, Eric Liang et al.ICLR 2020 · 17 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning ImplementationsRajdeep Singh Hundal, Yan Xiao, Xiaochun Cao, Jin Song Dong et al.ICSE 2025
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang et al.ICLR 2024 · 10 citations
- Investigating Multi-task Pretraining and Generalization in Reinforcement LearningAdrien Ali Taïga, Rishabh Agarwal, Jesse Farebrother, Aaron C. Courville et al.ICLR 2023
