Snake: A Variable-length Chain-based Prefetching for GPUs
Saba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
摘要
Graphics Processing Units (GPUs) utilize memory hierarchy and Thread-Level Parallelism (TLP) to tolerate off-chip memory latency, which is a significant bottleneck for memory-bound applications. However, parallel threads generate a large number of memory requests, which increases the average memory latency and degrades cache performance due to high contention. Prefetching is an effective technique to reduce memory access latency, and prior research shows the positive impact of stride-based prefetching on GPU performance. However, existing prefetching methods only rely on fixed strides. To address this limitation, this paper proposes a new prefetching technique, Snake, which is built upon chains of variable strides, using throttling and memory decoupling strategies. Snake achieves 80% coverage and 75% accuracy in prefetching demand memory requests, resulting in a 17% improvement in total GPU performance and energy consumption for memory-bound General-Purpose Graphics Processing Unit (GPGPU) applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- The Sparsity-Aware LazyGPU ArchitectureChangxi Liu, Miao Yu, Yifan Sun, Trevor E. CarlsonISCA 2025 · 被引用 1 次
- NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUsXia Zhao, Guangda Zhang, Lu Wang, Shiqing Zhang 等HPCA 2025 · 被引用 3 次
- SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUsJiwon Lee, Ju Min Lee, Yunho Oh, William J. Song 等HPCA 2023 · 被引用 22 次
- Morpheus: Extending the Last Level Cache Capacity in GPU Systems Using Idle GPU Core ResourcesSina Darabi, Mohammad Sadrosadati, Negar Akbarzadeh, Joël Lindegger 等MICRO 2022 · 被引用 24 次
- Analyzing and Leveraging Decoupled L1 Caches in GPUsMohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh 等HPCA 2021 · 被引用 30 次
