Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources
Jagadish B. Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. Loh
摘要
Many GPU applications issue irregular memory accesses to a very large memory footprint. We confirm observations from prior work that these irregular access patterns are severely bottlenecked by insufficient Translation Lookaside Buffer (TLB) reach, resulting in expensive page table walks. In this work, we investigate mechanisms to improve TLB reach without increasing the page size or the size of the TLB itself. Our work is based around the observation that a GPU’s instruction cache (I-cache) and Local Data Share (LDS) scratchpad memory are under-utilized in many applications, including those that suffer from poor TLB reach. We leverage this to opportunistically utilize idle capacity and port bandwidth from the GPU’s I-cache and LDS structures for address translations. We explore various potential architectural designs for each structure to optimize performance and minimize complexity. Both structures are organized as a victim cache between the L1 and L2 TLBs to boost translation reach. We find that our designs can increase performance on average by 30.1% without impacting the performance of applications that do not require additional reach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 被引用 21 次
- Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache ResourcesKonstantinos Kanellopoulos, Hong Chul Nam, Nisa Bostanci, Rahul Bera 等MICRO 2023 · 被引用 16 次
- A Case for Speculative Address Translation with Rapid Validation for GPUsJunhyeok Park, Osang Kwon, Yongho Lee, Seongwook Kim 等MICRO 2024 · 被引用 13 次
- Virtuoso: Enabling Fast and Accurate Virtual Memory Research via an Imitation-based Operating System Simulation MethodologyKonstantinos Kanellopoulos, Konstantinos Sgouras, F. Nisa Bostanci, Andreas Kosmas Kakolyris 等ASPLOS 2025 · 被引用 8 次
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 被引用 3 次
相关 Paper
- LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR CompressionYeonan Ha, Jiho Park, Hanna Cha, Jiwon Lee 等MICRO 2025 · 被引用 2 次
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 被引用 33 次
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 被引用 9 次
- SoftWalker: Supporting Software Page Table Walk for Irregular GPU ApplicationsSungbin Jang, Junhyeok Park, Yongho Lee, Osang Kwon 等MICRO 2025 · 被引用 4 次
- HIVE: A High-Priority Victim Cache for Accelerating GPU Memory AccessesYuhan Tang, Jianmin Zhang, Sheng Ma, Tiejun Li 等DAC 2025
