GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism
Yuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang, Tony Tong Geng, Ang Li, Yufei Ding
Abstract
With the increasing popularity of robotics in industrial control and autonomous driving, deep reinforcement learning (DRL) raises the attention of various fields. However, DRL computation on the modern powerful multi-GPU platform is still inefficient due to its heterogeneous tasks and complicated inter-task interactions. To this end, we propose GMI-DRL, the first systematic design for scaling multi-GPU DRL via adaptive-grained parallelism. To facilitate such a new parallelism scheme, GMI-DRL introduces a new concept -GPU Multiplexing Instance (GMI), a unified resource-adjustable sub-GPU design for heterogeneous tasks in DRL scaling. Besides, GMI-DRL introduces an adaptive Coordinator to effectively manage workloads and resources for better system performance. GMI-DRL also incorporates a specialized Communicator with highly efficient inter-GMI communication support to meet diverse communication demands. Extensive experiments demonstrate that GMI-DRL outperforms stateof-the-art DRL accelerating solution in training throughput (up to 2.34×) and GPU utilization (up to 40.8% improvement) on the DGX-A100 platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c47188da-4a08-4ef3-9216-d2950c476f4bBuilds on5
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 152 citations
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
- Accelerating Reinforcement Learning through GPU Atari EmulationSteven Dalton, Iuri FrosioNeurIPS 2020 · 50 citations
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang et al.ICLR 2020 · 32 citations
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen et al.USENIX ATC 2023 · 9 citations
Related papers
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Trainingzhixin wang, Jiaming Xu, Tianyi Zhou, Mingjun Zhang et al.ICML 2026 · 14 citations
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu et al.OSDI 2026
- Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesZhouyu He, Peng Qiao, Rongchun Li, Yong Dou et al.AAAI 2025 · 1 citation
- FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous PlatformsSamuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer et al.HPDC 2026 · 1 citation
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng et al.OSDI 2023 · 46 citations
