Lune

ISCA2026Top-tier venue

Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory Management

Xiangyue Huang, Yanan Guo, Yuanchao Xu

2026Year

Abstract

Multi-GPU systems have become popular to meet the growing demands for high parallelism and large memory capacity via unified virtual memory (UVM). However, performance is often constrained by non-uniform memory access (NUMA) overheads due to frequent data sharing across GPUs. Prior work adopts fine-grained page migration and duplication to reduce remote access overheads, but our characterization of recent NVLinks shows that such designs fail to fully exploit their capabilities. In particular, nonlinear latency-size scaling, negligible contention, and abundant bandwidth favor coarsegrained transfers. While coarse-grained approaches better utilize NVLink bandwidth, they can introduce excessive remote accesses and update overheads. We propose CDFD, a duplication-centric mechanism that combines coarse-grained duplication to maximize bandwidth utilization with selective fine-grained deduplication to mitigate unnecessary remote updates. By leveraging idle GPU memory capacity and dynamically refining duplication decisions, CDFD balances performance and overhead. Experimental results show that CDFD achieves average performance improvements of 66% and 65% over state-of-the-art methods GPS and GRIT, respectively.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get d65a1ee5-db07-4715-8302-530814135511

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines