Lune

ISCA2025顶会

Forest: Access-aware GPU UVM Management

Mao Lin, Yuan Feng, Guilherme Cox, Hyeran Jeon

2025年份
9被引次数
1顶会引用

摘要

With GPU unified virtual memory (UVM), CPU and GPU can share a flat virtual address space. UVM enables the GPUs to utilize the larger CPU system memory as an expanded memory space. However, UVM’s on-demand page migration is accompanied by expensive page fault handling overhead. To mitigate such overhead, tree-based neighboring prefetcher (TBNp) has been used by GPUs. TBNp effectively reduces page faults by exploiting locality at multiple levels. However, we observe its access-pattern oblivious design leads to excessive page thrashing and unnecessary migrations. In this paper, we tackle the inefficiencies with a novel access-aware UVM management, Forest. Forest uses a software-hardware codesign to configure the optimal tree prefetchers at runtime based on each data object’s access patterns. With the heterogeneous tree-based prefetching, Forest provides 1.86 × and 1.39 × speedups over the baseline TBNp and state-of-the-art optimization solutions, respectively. Forest also shows a 1.51 × speedup for real-world deep learning models, including CNNs and Transformers.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖