SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
Zhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, Yi Wu
摘要
The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLly Scalable RL (SRL), which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wallclock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL source code is available at: https://github.com/openpsi-project/srl .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Hao Wang, Devesh Tiwari, Jian Li 等SC 2024 · 被引用 10 次
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang 等EuroSys 2026 · 被引用 2 次
- Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesZhouyu He, Peng Qiao, Rongchun Li, Yong Dou 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper9
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
- Towards Playing Full MOBA Games with Deep Reinforcement LearningDeheng Ye, Guibin Chen, Wen Zhang, Sheng Chen 等NeurIPS 2020 · 被引用 225 次
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme 等ICML 2020 · 被引用 131 次
相关 Paper
- RLlib Flow: Distributed Reinforcement Learning is a Dataflow ProblemEric Liang, Zhanghao Wu, Michael Luo, Sven Mika 等NeurIPS 2021 · 被引用 39 次
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Trainingzhixin wang, Jiaming Xu, Tianyi Zhou, Mingjun Zhang 等ICML 2026 · 被引用 14 次
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu 等OSDI 2026
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen 等USENIX ATC 2023 · 被引用 9 次
- An Extensible, Data-Oriented Architecture for High-Performance, Many-World SimulationBrennan Shacklett, Luc Guy Rosenzweig, Zhiqiang Xie, Bidipta Sarkar 等SIGGRAPH 2023 · 被引用 13 次
