In-depth analyses of unified virtual memory system for GPU accelerated computing
Tyler N. Allen, Rong Ge
摘要
The abstraction of a shared memory space over separate CPU and GPU memory domains has eased the burden of portability for many HPC codebases. However, users pay for the ease of use provided by systems-managed memory space with a moderate-to-high performance overhead. NVIDIA Unified Virtual Memory (UVM) is presently the primary real-world implementation of such abstraction and offers a functionally equivalent testbed for a novel in-depth performance study for both UVM and future Linux Heterogeneous Memory Management (HMM) compatible systems. The continued advocation for UVM and HMM motivates the improvement of the underlying system. We focus on a UVM-based system and investigate the root causes of the UVM overhead, which is a non-trivial task due to the complex interactions of multiple hardware and software constituents and the requirement of targeted analysis methodology.
In this paper, we take a deep dive into the UVM system architecture and the internal behaviors of page fault generation and servicing. We reveal specific GPU hardware limitations using targeted benchmarks to uncover driver functionality as a real-time system when processing the resultant workload. We further provide a quantitative evaluation of fault handling for various applications under different scenarios, including prefetching and oversubscription. We find that the driver workload is dependent on the interactions among application access patterns, GPU hardware constraints, and Host OS components. We determine that the cost of host OS components is significant and present across implementations, warranting close attention. This study serves as a proxy for future shared memory systems such as those that interface with HMM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementWonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong SimOSDI 2024 · 被引用 248 次
- SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMMGerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao 等ISCA 2023 · 被引用 27 次
- Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryJeongmin Hong, Sungjun Cho, Geonwoo Park, Wonhyuk Yang 等HPCA 2024 · 被引用 21 次
- G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor MigrationsHaoyang Zhang, Yirui Eric Zhou, Yuqi Xue, Yiqi Liu 等MICRO 2023 · 被引用 21 次
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang 等HPCA 2024 · 被引用 19 次
它引用的顶会 Paper4
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
- EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUsSeungwon Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong 等VLDB 2021 · 被引用 66 次
- Traversing Large Graphs on GPUs with Unified MemoryPrasun Gera, Hyojong Kim, Piyush Sao, Hyesoon Kim 等VLDB 2020 · 被引用 58 次
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder 等HPCA 2020 · 被引用 50 次
相关 Paper
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 被引用 7 次
- Observability-Aided Gpu Memory OversubscriptionPratheek B, Khushit Shah, Arkaprava BasuISCA 2026
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 被引用 2 次
- HELM: Characterizing Unified Memory Accesses to Improve GPU Performance under Memory OversubscriptionNathan Jones, Tyler N. Allen, Rong GeSC 2025 · 被引用 5 次
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 被引用 9 次
