IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE Invalidations
Bingyao Li, Yanan Guo, Yueqi Wang, Aamer Jaleel, Jun Yang, Xulong Tang
Abstract
Multi-GPU systems have emerged as a desirable platform to deliver high computing capabilities and large memory capacity to accommodate large dataset sizes. However, naively employing multi-GPU incurs non-scalable performance. One major reason is that execution efficiency suffers expensive address translations in multi-GPU systems. The data-sharing nature of GPU applications requires page migration between GPUs to mitigate non-uniform memory access overheads. Unfortunately, frequent page migration incurs substantial page table invalidation overheads to ensure translation coherence. A comprehensive investigation of multi-GPU address translation efficiency identifies two significant bottlenecks caused by page table invalidation requests: (i) increased latency for demand TLB miss requests and (ii) increased waiting latency for performing page migrations. Based on observations, we propose IDYLL, which reduces the number of page table invalidations by maintaining an “in-PTE" directory and reduces invalidation latency by batching multiple invalidation requests to exploit spatial locality. We show that IDYLL improves overall performance by 69.9% on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e12dabf2-242b-44e1-aee7-0614ac747e72Cited by top-tier papers6
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 20 citations
- GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementYueqi Wang, Bingyao Li, Aamer Jaleel, Jun Yang et al.HPCA 2024 · 19 citations
- A Case for Speculative Address Translation with Rapid Validation for GPUsJunhyeok Park, Osang Kwon, Yongho Lee, Seongwook Kim et al.MICRO 2024 · 13 citations
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout et al.HPCA 2025 · 12 citations
- STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUBingyao Li, Yueqi Wang, Tianyu Wang, Lieven Eeckhout et al.MICRO 2024 · 11 citations
Builds on22
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 73 citations
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 71 citations
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe et al.ASPLOS 2020 · 62 citations
- Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for ParallelismDimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, Josep TorrellasASPLOS 2020 · 55 citations
- Griffin: Hardware-Software Support for Efficient Page Migration in Multi-GPU SystemsTrinayan Baruah, Yifan Sun, Ali Tolga Dinçer, Saiful A. Mojumder et al.HPCA 2020 · 50 citations
Related papers
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingBingyao Li, Jieming Yin, Anup Holey, Youtao Zhang et al.HPCA 2023 · 25 citations
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee et al.MICRO 2025 · 2 citations
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU SystemsJunsung Kim, Dongho Ha, Sungwoo Kim, Wonho Cho et al.ISCA 2026
