Lune

MICRO2025顶会

LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression

Yeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee, Joonsung Kim, Won Woo Ro, Youngsok Kim

2025年份
2被引次数
1顶会引用

摘要

Modern Graphics Processing Units (GPUs) support virtual memory to ease programmability and concurrency, but still suffer from significant address translation overhead due to frequent Translation Lookaside Buffer (TLB) misses and limited TLB Miss-Status Holding Register (MSHR) capacity.These misses trigger long-latency page table walks and stall memory accesses, degrading overall performance.In this paper, we present LATPC, a novel mechanism that combines TLB prefetching with MSHR compression to accelerate GPU address translation.LATPC leverages the regularity of Virtual Page Numbers (VPNs) across threads within a warp to coalesce TLB misses at the warp instruction level.LATPC then compresses multiple TLB miss requests into a small number of MSHR entries, reducing contention and enabling more efficient resource use.LATPC further improves translation efficiency by batching page table walks based on the identified VPN patterns, which not only reduces the number of page table walk invocations but also increases off-chip DRAM row buffer locality, lowering the latency of memory accesses during translation.Our evaluation using 24 GPU workloads shows that LATPC effectively exploits regularities and localities in address translation requests within a warp, achieving a 1.47× geometric mean speedup over the baseline without TLB prefetching.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖