LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression
Yeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee, Joonsung Kim, Won Woo Ro, Youngsok Kim
Abstract
Modern Graphics Processing Units (GPUs) support virtual memory to ease programmability and concurrency, but still suffer from significant address translation overhead due to frequent Translation Lookaside Buffer (TLB) misses and limited TLB Miss-Status Holding Register (MSHR) capacity.These misses trigger long-latency page table walks and stall memory accesses, degrading overall performance.In this paper, we present LATPC, a novel mechanism that combines TLB prefetching with MSHR compression to accelerate GPU address translation.LATPC leverages the regularity of Virtual Page Numbers (VPNs) across threads within a warp to coalesce TLB misses at the warp instruction level.LATPC then compresses multiple TLB miss requests into a small number of MSHR entries, reducing contention and enabling more efficient resource use.LATPC further improves translation efficiency by batching page table walks based on the identified VPN patterns, which not only reduces the number of page table walk invocations but also increases off-chip DRAM row buffer locality, lowering the latency of memory accesses during translation.Our evaluation using 24 GPU workloads shows that LATPC effectively exploits regularities and localities in address translation requests within a warp, achieving a 1.47× geometric mean speedup over the baseline without TLB prefetching.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- SoftWalker: Supporting Software Page Table Walk for Irregular GPU ApplicationsSungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon et al.MICRO 2025 · 4 citations
- Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputJiwon Lee, Gun Ko, Myung Kuk Yoon, Ipoom Jeong et al.HPCA 2025 · 4 citations
- Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip ResourcesJagadish B. Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. LohMICRO 2021 · 17 citations
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUsJiwon Lee, Ju Min Lee, Yunho Oh, William J. Song et al.HPCA 2023 · 22 citations
