Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU Systems
Junsung Kim, Dongho Ha, Sungwoo Kim, Wonho Cho, Sungbin Kim, Yufei Ding, Won Woo Ro
摘要
Unified Virtual Memory in multi-GPU systems provides scalable hardware resources with simplified memory management to users. However, its performance is often restricted by high-latency page faults. When a page migrates, the migration mechanism invalidates old mappings on all GPUs but updates the new mapping only on the destination GPU. This leaves other GPUs' page table entries invalid, generating subsequent page faults when the non-destination GPUs attempt to access the migrated page. Additionally, we observe that naively updating all GPUs' entries incurs extra page table walks, resulting in significant latency overhead. To address this, we propose ShadowUpdate, a redesign of the migration handling mechanism to reduce page faults by leveraging the invalidation phase. ShadowUpdate exploits the existing invalidation broadcast to proactively propagate the new mapping simultaneously. By combining invalidation and mapping updates, this approach eliminates redundant page faults and the associated page table walks, accelerating the overall migration process. To ensure correctness, ShadowUpdate also includes a lightweight in-flight migration tracker that holds translation requests during migration, records the new mapping at invalidation, and releases the requests once the page copy completes to prevent invalid accesses. ShadowUpdate improves overall performance by on average over a baseline UVM design across 14 representative multi-GPU UVM workloads.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingBingyao Li, Jieming Yin, Anup Holey, Youtao Zhang 等HPCA 2023 · 被引用 25 次
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 被引用 9 次
- Coarse-Grained Duplication First, Fine-Grained Deduplication Later: Duplication-Centric Multi-GPU Memory ManagementXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 被引用 73 次
