Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB Design
Bingyao Li, Jieming Yin, Youtao Zhang, Xulong Tang
Abstract
In recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8ca5b44-1963-4472-9e5a-6c6be6a37043Cited by top-tier papers10
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 21 citations
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 20 citations
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang et al.HPCA 2024 · 19 citations
- IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsBingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel et al.MICRO 2023 · 16 citations
- A Case for Speculative Address Translation with Rapid Validation for GPUsJunhyeok Park, Osang Kwon, Yongho Lee, Seongwook Kim et al.MICRO 2024 · 13 citations
Builds on4
- Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUsEsha Choukse, Michael B. Sullivan, Mike O'Connor, Mattan Erez et al.ISCA 2020 · 58 citations
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
- NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing UnitsBongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim et al.ASPLOS 2020 · 29 citations
- Improving GPU Multi-tenancy with Page Walk StealingB Pratheek, Neha Jawalkar, Arkaprava BasuHPCA 2021 · 26 citations
Related papers
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee et al.MICRO 2025 · 2 citations
- STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUBingyao Li, Yueqi Wang, Tianyu Wang, Lieven Eeckhout et al.MICRO 2024 · 11 citations
- SoftWalker: Supporting Software Page Table Walk for Irregular GPU ApplicationsSungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon et al.MICRO 2025 · 4 citations
- Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip ResourcesJagadish B. Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. LohMICRO 2021 · 17 citations
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 9 citations
