Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory Management
Xiangyue Huang, Yanan Guo, Yuanchao Xu
摘要
Multi-GPU systems have become popular to meet the growing demands for high parallelism and large memory capacity via unified virtual memory (UVM). However, performance is often constrained by non-uniform memory access (NUMA) overheads due to frequent data sharing across GPUs. Prior work adopts fine-grained page migration and duplication to reduce remote access overheads, but our characterization of recent NVLinks shows that such designs fail to fully exploit their capabilities. In particular, nonlinear latency-size scaling, negligible contention, and abundant bandwidth favor coarsegrained transfers. While coarse-grained approaches better utilize NVLink bandwidth, they can introduce excessive remote accesses and update overheads. We propose CDFD, a duplication-centric mechanism that combines coarse-grained duplication to maximize bandwidth utilization with selective fine-grained deduplication to mitigate unnecessary remote updates. By leveraging idle GPU memory capacity and dynamically refining duplication decisions, CDFD balances performance and overhead. Experimental results show that CDFD achieves average performance improvements of 66% and 65% over state-of-the-art methods GPS and GRIT, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang 等HPCA 2024 · 被引用 19 次
- Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU SystemsJunsung Kim, Dongho Ha, Sungwoo Kim, Wonho Cho 等ISCA 2026
- LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page PrefetcherXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout 等HPCA 2025 · 被引用 12 次
- GPS: A Global Publish-Subscribe Model for Multi-GPU Memory ManagementHarini Muthukrishnan, Daniel Lustig, David W. Nellans, Thomas F. WenischMICRO 2021 · 被引用 23 次
