Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained Transfers
Harini Muthukrishnan, David W. Nellans, Daniel Lustig, Jeffrey A. Fessler, Thomas F. Wenisch
Abstract
Despite continuing research into inter-GPU communication mechanisms, extracting performance from multi-GPU systems remains a significant challenge. Inter-GPU communication via bulk DMA-based transfers exposes data transfer latency on the GPU’s critical execution path because these large transfers are logically interleaved between compute kernels. Conversely, fine-grained peer-to-peer memory accesses during kernel execution lead to memory stalls that can exceed the GPUs’ ability to cover these operations via multi-threading. Worse yet, these sub-cacheline transfers are highly inefficient on current inter-GPU interconnects. To remedy these issues, we propose PROACT, a system enabling remote memory transfers with the programmability and pipeline advantages of peer-to-peer stores, while achieving interconnect efficiency that rivals bulk DMA transfers. Combining compile-time instrumentation with fine-grain tracking of data block readiness within each GPU, PROACT enables interconnect-friendly data transfers while hiding the transfer latency via pipelining during kernel execution. This work describes both hardware and software implementations of PROACT and demonstrates the effectiveness of a PROACT software prototype on three generations of GPU hardware and interconnects. Achieving near-ideal interconnect efficiency, PROACT realizes a mean speedup of 3.0× over single-GPU performance for 4-GPU systems, capturing 83% of available performance opportunity. On a 16-GPU NVIDIA DGX-2 system, we demonstrate an 11.0× average strong-scaling speedup over single-GPU performance, 5.3× better than a bulk DMA-based approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fbd3d23-d268-433f-9cb4-375e2a91e035Cited by top-tier papers5
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang et al.MobiCom 2024 · 40 citations
- GPS: A Global Publish-Subscribe Model for Multi-GPU Memory ManagementHarini Muthukrishnan, Daniel Lustig, David W. Nellans, Thomas F. WenischMICRO 2021 · 23 citations
- FinePack: Transparently Improving the Efficiency of Fine-Grained Transfers in Multi-GPU SystemsHarini Muthukrishnan, Daniel Lustig, Oreste Villa, Thomas F. Wenisch et al.HPCA 2023 · 14 citations
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang et al.USENIX ATC 2025 · 1 citation
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang et al.ISCA 2026
Builds on3
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel et al.HPCA 2020 · 38 citations
Related papers
- ARK: GPU-driven Code Execution for Distributed Deep LearningChangho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu et al.NSDI 2023 · 22 citations
- RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAsMarco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer et al.EuroSys 2026 · 1 citation
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He et al.VLDB 2026
- Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony BuffersJianqin Yan, Shi Qiu, Yina Lv, Yifan Hu et al.SC 2025 · 3 citations
- A Simple Cache Coherence Scheme for Integrated CPU-GPU SystemsArdhi Wiratama Baskara Yudha, Reza Pulungan, Henry Hoffmann, Yan SolihinDAC 2020 · 3 citations
