Forest: Access-aware GPU UVM Management
Mao Lin, Yuan Feng, Guilherme Cox, Hyeran Jeon
Abstract
With GPU unified virtual memory (UVM), CPU and GPU can share a flat virtual address space. UVM enables the GPUs to utilize the larger CPU system memory as an expanded memory space. However, UVM’s on-demand page migration is accompanied by expensive page fault handling overhead. To mitigate such overhead, tree-based neighboring prefetcher (TBNp) has been used by GPUs. TBNp effectively reduces page faults by exploiting locality at multiple levels. However, we observe its access-pattern oblivious design leads to excessive page thrashing and unnecessary migrations. In this paper, we tackle the inefficiencies with a novel access-aware UVM management, Forest. Forest uses a software-hardware codesign to configure the optimal tree prefetchers at runtime based on each data object’s access patterns. With the heterogeneous tree-based prefetching, Forest provides 1.86 × and 1.39 × speedups over the baseline TBNp and state-of-the-art optimization solutions, respectively. Forest also shows a 1.51 × speedup for real-world deep learning models, including CNNs and Transformers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 59931d2f-641e-47bd-a65c-883c3af1e47fCited by top-tier papers1
Ask how each one uses itRelated papers
- LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page PrefetcherXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- SUV: Static Analysis Guided Unified Virtual MemoryPratheek B, Guilherme Cox, Ján Veselý, Arkaprava BasuMICRO 2024 · 7 citations
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 9 citations
- ARIADNE: Adaptive UVM Management for Efficient GPU Memory OversubscriptionHyunkyun Shin, Seongtae Bang, Hyungwon Park, Daehoon KimHPCA 2026 · 2 citations
- DeepUM: Tensor Migration and Prefetching in Unified MemoryJaehoon Jung, Jinpyo Kim, Jaejin LeeASPLOS 2023 · 33 citations
