Improving GPU Multi-tenancy with Page Walk Stealing
B Pratheek, Neha Jawalkar, Arkaprava Basu
Abstract
GPU (Graphics Processing Unit) architecture has evolved to accelerate parts of a single application at a time. Consequently, several aspects of its architecture, particularly the virtual memory, have embraced a shared-mostly design. This implicitly assumes that a single application and, thus, one address space is resident in the GPU at a time. However, recent trends, e.g., deployment of GPUs in the cloud, necessitate efficient multi-tenancy. Multi-tenancy is needed for sharing the physical resources of a large server-class GPU across multiple concurrent tenants (applications) for resource consolidation while ensuring fairness among the tenants.
We first quantify how different components of GPU's virtual memory can impede multi-tenancy. We show that shared page walkers are a key bottleneck under multi-tenancy. We, therefore, propose dynamic page walk stealing that enables soft partitioning of the shared pool of walkers -reducing destructive interference between the tenants while also aggregating resources where possible. Over today's design, we improve throughput by 37%, and weighted IPC by 15%, on average, over 45 workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c9de6ec-c230-4912-8240-0c8897e1f68cCited by top-tier papers11
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 21 citations
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 20 citations
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang et al.HPCA 2024 · 19 citations
Related papers
- Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputJiwon Lee, Gun Ko, Myung Kuk Yoon, Ipoom Jeong et al.HPCA 2025 · 4 citations
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- SoftWalker: Supporting Software Page Table Walk for Irregular GPU ApplicationsSungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon et al.MICRO 2025 · 4 citations
- UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource EfficiencyXia Zhao, Guangda Zhang, Lu Wang, Huadong DaiISCA 2025 · 1 citation
- Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingBingyao Li, Jieming Yin, Anup Holey, Youtao Zhang et al.HPCA 2023 · 25 citations
