Highly Parallelized Reinforcement Learning Training with Relaxed Assignment Dependencies
Zhouyu He, Peng Qiao, Rongchun Li, Yong Dou, Yusong Tan
Abstract
As the demands for superior agents grow, the training complexity of Deep Reinforcement Learning (DRL) becomes higher. Thus, accelerating training of DRL has become a major research focus. Dividing the DRL training process into subtasks and using parallel computation can effectively reduce training costs. However, current DRL training systems lack sufficient parallelization due to data assignment between subtask components. This assignment issue has been ignored, but addressing it can further boost training efficiency. Therefore, we propose a high-throughput distributed RL training system called TianJi. It relaxes assignment dependencies between subtask components and enables event-driven asynchronous communication. Meanwhile, TianJi maintains clear boundaries between subtask components. To address convergence uncertainty from relaxed assignment dependencies, TianJi proposes a distributed strategy based on the balance of sample production and consumption. The strategy controls the staleness of samples to correct their quality, ensuring convergence. We conducted extensive experiments. TianJi achieves a convergence time acceleration ratio of up to 4.37 compared to related comparison systems. When scaled to eight computational nodes, TianJi shows a convergence time speedup of 1.6 and a throughput speedup of 7.13 relative to XingTian, demonstrating its capability to accelerate training and scalability. In data transmission efficiency experiments, TianJi significantly outperforms other systems, approaching hardware limits. TianJi also shows effectiveness in on-policy algorithms, achieving convergence time acceleration ratios of 4.36 and 2.95 compared to RLlib and XingTian. TianJi is accessible at https://github.com/HiPRL/TianJi.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca96fdb7-c514-4022-bf97-d86762e4f9d1Builds on10
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme et al.ICML 2020 · 131 citations
- Large Language Models as Generalizable Policies for Embodied TasksAndrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure et al.ICLR 2024 · 114 citations
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu et al.NeurIPS 2023 · 34 citations
- Heterogeneous Skill Learning for Multi-agent TasksYuntao Liu, Yuan Li, Xinhai Xu, Yong Dou et al.NeurIPS 2022 · 33 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
Related papers
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang et al.ICLR 2024 · 10 citations
- Stellaris: Staleness-Aware Distributed Reinforcement Learning with Serverless ComputingHanfei Yu, Hao Wang, Devesh Tiwari, Jian Li et al.SC 2024 · 10 citations
- GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained ParallelismYuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang et al.USENIX ATC 2025 · 3 citations
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target NetworksMichael Luo, Jiahao Yao, Richard Liaw, Eric Liang et al.ICLR 2020 · 17 citations
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen et al.USENIX ATC 2023 · 9 citations
