SoftWalker: Supporting Software Page Table Walk for Irregular GPU Applications
Sungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon, Donghyun Kim, Juyoung Seok, Seokin Hong
Abstract
Address translation has become a significant and growing performance bottleneck in modern GPUs, especially for emerging irregular applications with high TLB miss rates.The limited concurrency of hardware Page Table Walkers (PTWs), due to their small and fixed number, causes severe contention and substantial queueing delays under high translation pressure, which significantly degrades performance.This paper introduces SoftWalker, a novel, scalable, and flexible framework that fundamentally shifts the GPU page table walking from fixed-function hardware to software execution.SoftWalker leverages the GPU's massive thread-level parallelism by dynamically dispatching specialized, lightweight software threads running on GPU cores to handle TLB misses requiring page table walks.In addition, to expand L2 TLB MSHR capacity on demand, SoftWalker incorporates In-TLB MSHRs, a key innovation that repurposes underutilized L2 TLB entries to track outstanding misses when existing MSHRs are saturated.By alleviating MSHR-induced contention, this design preserves the key advantage of highly parallel page table walking in software.SoftWalker enables thousands of concurrent page table walks, significantly reducing PTW-level contention and translation queueing delays.As a result, it achieves an average reduction of 72.8% in page walk latency and delivers an average speedup of 2.24× (3.94× for irregular workloads).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 94a68088-961c-4646-8a72-af6a71b2893dCited by top-tier papers1
Ask how each one uses itRelated papers
- Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputJiwon Lee, Gun Ko, Myung Kuk Yoon, Ipoom Jeong et al.HPCA 2025 · 4 citations
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee et al.MICRO 2025 · 2 citations
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip ResourcesJagadish B. Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. LohMICRO 2021 · 17 citations
- SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUsJiwon Lee, Ju Min Lee, Yunho Oh, William J. Song et al.HPCA 2023 · 22 citations
