USENIX ATC2025顶会
GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism
Yuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang, Tony Tong Geng, Ang Li, Yufei Ding
摘要
With the increasing popularity of robotics in industrial control and autonomous driving, deep reinforcement learning (DRL) raises the attention of various fields. However, DRL computation on the modern powerful multi-GPU platform is still inefficient due to its heterogeneous tasks and complicated inter-task interactions. To this end, we propose GMI-DRL, the first systematic design for scaling multi-GPU DRL via adaptive-grained parallelism. To facilitate such a new parallelism scheme, GMI-DRL introduces a new concept -GPU Multiplexing Instance (GMI), a unified resource-adjustable sub-GPU design for heterogeneous tasks in DRL scaling. Besides, GMI-DRL introduces an adaptive Coordinator to effectively manage workloads and resources for better system performance. GMI-DRL also incorporates a specialized Communicator with highly efficient inter-GMI communication support to meet diverse communication demands. Extensive experiments demonstrate that GMI-DRL outperforms stateof-the-art DRL accelerating solution in training throughput (up to 2.34×) and GPU utilization (up to 40.8% improvement) on the DGX-A100 platform.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 被引用 152 次
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 被引用 91 次
- Accelerating Reinforcement Learning through GPU Atari EmulationSteven Dalton, Iuri FrosioNeurIPS 2020 · 被引用 50 次
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang 等ICLR 2020 · 被引用 32 次
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen 等USENIX ATC 2023 · 被引用 9 次
相关 Paper
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Trainingzhixin wang, Jiaming Xu, Tianyi Zhou, Mingjun Zhang 等ICML 2026 · 被引用 14 次
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu 等OSDI 2026
- Highly Parallelized Reinforcement Learning Training with Relaxed Assignment DependenciesZhouyu He, Peng Qiao, Rongchun Li, Yong Dou 等AAAI 2025 · 被引用 1 次
- FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous PlatformsSamuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer 等HPDC 2026 · 被引用 1 次
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
