Lune

MICRO2025Top-tier venue

LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression

Yeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee, Joonsung Kim, Won Woo Ro, Youngsok Kim

2025Year
2Citations
1Top-tier citations

Abstract

Modern Graphics Processing Units (GPUs) support virtual memory to ease programmability and concurrency, but still suffer from significant address translation overhead due to frequent Translation Lookaside Buffer (TLB) misses and limited TLB Miss-Status Holding Register (MSHR) capacity.These misses trigger long-latency page table walks and stall memory accesses, degrading overall performance.In this paper, we present LATPC, a novel mechanism that combines TLB prefetching with MSHR compression to accelerate GPU address translation.LATPC leverages the regularity of Virtual Page Numbers (VPNs) across threads within a warp to coalesce TLB misses at the warp instruction level.LATPC then compresses multiple TLB miss requests into a small number of MSHR entries, reducing contention and enabling more efficient resource use.LATPC further improves translation efficiency by batching page table walks based on the identified VPN patterns, which not only reduces the number of page table walk invocations but also increases off-chip DRAM row buffer locality, lowering the latency of memory accesses during translation.Our evaluation using 24 GPU workloads shows that LATPC effectively exploits regularities and localities in address translation requests within a warp, achieving a 1.47× geometric mean speedup over the baseline without TLB prefetching.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines