ARK: GPU-driven Code Execution for Distributed Deep Learning
Changho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu, Peng Cheng, Yongqiang Xiong
摘要
Modern state-of-the-art deep learning (DL) applications tend to scale out to a large number of parallel GPUs. Unfortunately, we observe that the collective communication overhead across GPUs is often the key limiting factor of performance for distributed DL. It under-utilizes the networking bandwidth by frequent transfers of small data chunks, which also incurs a substantial I/O overhead on GPU that interferes with computation on GPU. The root cause lies in the inefficiency of CPU-based communication event handling as well as the inability to control the GPU's internal DMA engine with GPU threads.
To address the problem, we propose a GPU-driven code execution system that leverages a GPU-controlled hardware DMA engine for I/O offloading. Our custom DMA engine pipelines multiple DMA requests to support efficient small data transfer while it eliminates the I/O overhead on GPU cores. Unlike existing GPU DMA engines initiated only by CPU, we let GPU threads directly control DMA operations, which leads to a highly efficient system where GPUs drive their own execution flow and handle communication events autonomously without CPU intervention. Our prototype DMA engine achieves a line-rate from a message size as small as 8KB (3.9x better throughput) with only 4.3µs of communication latency (9.1x faster) while it incurs little interference with computation on GPU, achieving 1.8x higher all-reduce throughput in a real training workload.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An 等OSDI 2026 · 被引用 40 次
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin 等MICRO 2024 · 被引用 26 次
- T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & CollectivesSuchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena 等ASPLOS 2024 · 被引用 13 次
- Optimizing Distributed ML Communication with Fused Computation-Collective OperationsKishore Punniyamurthy, Khaled Hamidouche, Bradford M. BeckmannSC 2024 · 被引用 11 次
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 被引用 102 次
- Enabling Compute-Communication Overlap in Distributed Deep Learning Training PlatformsSaeed Rashidi, Matthew Denton, Srinivas Sridharan, Sudarshan Srinivasan 等ISCA 2021 · 被引用 39 次
相关 Paper
- MSCCLang: Microsoft Collective Communication LanguageMeghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi 等ASPLOS 2023 · 被引用 40 次
- Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained TransfersHarini Muthukrishnan, David W. Nellans, Daniel Lustig, Jeffrey A. Fessler 等ISCA 2021 · 被引用 16 次
- RoCC: Harnessing Raster Operations Pipeline for Efficient Tensor Collective CommunicationYuan Feng, Daniel Wong, Hyeran JeonISCA 2026
- LAIKA: Machine Learning-Assisted In-Kernel APU AccelerationHaoming Zhuo, Dingding Li, Ronghua Lin, Yong TangASPLOS 2026
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 被引用 9 次
