GPS: A Global Publish-Subscribe Model for Multi-GPU Memory Management
Harini Muthukrishnan, Daniel Lustig, David W. Nellans, Thomas F. Wenisch
Abstract
Suboptimal management of memory and bandwidth is one of the primary causes of low performance on systems comprising multiple GPUs. Existing memory management solutions like Unified Memory (UM) offer simplified programming but come at the cost of performance: applications can even exhibit slowdown with increasing GPU count due to their inability to leverage system resources effectively. To solve this challenge, we propose GPS, a HW/SW multi-GPU memory management technique that efficiently orchestrates inter-GPU communication using proactive data transfers. GPS offers the programmability advantage of multi-GPU shared memory with the performance of GPU-local memory. To enable this, GPS automatically tracks the data accesses performed by each GPU, maintains duplicate physical replicas of shared regions in each GPU’s local memory, and pushes updates to the replicas in all consumer GPUs. GPS is compatible within the existing NVIDIA GPU memory consistency model but takes full advantage of its relaxed nature to deliver high performance. We evaluate GPS in the context of a 4-GPU system with varying interconnects and show that GPS achieves an average speedup of 3.0 × relative to the performance of a single GPU, outperforming the next best available multi-GPU memory management technique by 2.3 × on average. In a 16-GPU system, using a future PCIe 6.0 interconnect, we demonstrate a 7.9 × average strong scaling speedup over single-GPU performance, capturing 80% of the available opportunity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3271032c-4078-41e0-8334-01b9bbb5a73cCited by top-tier papers7
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang et al.HPCA 2024 · 19 citations
- SAC: Sharing-Aware Caching in Multi-Chip GPUsShiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven EeckhoutISCA 2023 · 18 citations
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel et al.MICRO 2023 · 16 citations
- FinePack: Transparently Improving the Efficiency of Fine-Grained Transfers in Multi-GPU SystemsHarini Muthukrishnan, Daniel Lustig, Oreste Villa, Thomas F. Wenisch et al.HPCA 2023 · 14 citations
- T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & CollectivesSuchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena et al.ASPLOS 2024 · 13 citations
Builds on4
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel et al.HPCA 2020 · 38 citations
- Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained TransfersHarini Muthukrishnan, David W. Nellans, Daniel Lustig, Jeffrey A. Fessler et al.ISCA 2021 · 16 citations
Related papers
- Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual MemorySangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun et al.USENIX ATC 2022 · 28 citations
- Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory ManagementXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 2 citations
- ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data AnalysisJie Zhang, Myoungsoo JungISCA 2020 · 14 citations
- Ohm-GPU: Integrating New Optical Network and Heterogeneous Memory into GPU Multi-ProcessorsJie Zhang, Myoungsoo JungMICRO 2021 · 6 citations
