Observability-Aided Gpu Memory Oversubscription
Pratheek B, Khushit Shah, Arkaprava Basu
摘要
Unified Virtual Memory (UVM) enables the oversubscription of GPUs' limited High Bandwidth Memory (HBM) capacity. Unfortunately, applications can significantly slow down when HBM is oversubscribed. We set out to make oversubscription practical without needing custom hardware modifications. The UVM driver, executing on the CPU, is responsible for eviction and prefetching but lacks observability into GPUs' accesses to HBM-resident memory. This fundamentally limits the driver's ability to make informed decisions. Toward this, we repurpose existing hardware access counters originally designed to track PCIe accesses to CPU's DRAM to aid page migration. Instead, we leverage the counters to provide (sampled) observability into GPU's accesses to HBM-resident pages. We then create ObservUVM, a novel software framework that enables easy exploration of observability-aided custom eviction and prefetching policies in the userspace, facilitating future research. We demonstrate that better-informed eviction and prefetching policies, enabled by the newfound observability, can significantly speed up GPU applications under memory oversubscription.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 被引用 7 次
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 被引用 2 次
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 被引用 73 次
- HELM: Characterizing Unified Memory Accesses to Improve GPU Performance under Memory OversubscriptionNathan Jones, Tyler N. Allen, Rong GeSC 2025 · 被引用 5 次
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 被引用 9 次
