R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUs
Dongho Ha, Yunho Oh, Won Woo Ro
Abstract
A generally used GPU programming methodology is that adjacent threads access data in neighbor or specific-stride memory addresses and perform computations with the fetched data. This paper demonstrates that the memory addresses often exhibit a simple linear value pattern across GPU threads, as each thread uses built-in variables and constant values to compute the memory addresses. However, since the threads compute their context data individually, GPUs incur a heavy instruction overhead to calculate the memory addresses, even though they exhibit a simple pattern. We propose a GPU architecture called Removing ReDunDancy Utilizing Linearity of Address Generation (R2D2), reducing a large amount of the dynamic instruction count by detecting such linear patterns in the memory addresses and exploiting them for kernel computations. R2D2 detects linearities of the memory addresses with software support and pre-computes them before the threads execute the instructions. With the proposed scheme, each thread is able to compute its memory addresses with fewer dynamic instructions than conventional GPUs. In our evaluation, R2D2 achieves dynamic instruction reduction by 28%, 1.25x speedup, and energy consumption reduction by 17% over baseline GPU.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin et al.MICRO 2024 · 26 citations
- Photon: A Fine-grained Sampled Simulation Methodology for GPU WorkloadsChangxi Liu, Yifan Sun, Trevor E. CarlsonMICRO 2023 · 10 citations
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 3 citations
Related papers
- Dimensionality-Aware Redundant SIMT Instruction EliminationTsung Tai Yeh, Roland N. Green, Timothy G. RogersASPLOS 2020 · 12 citations
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim et al.MICRO 2020 · 27 citations
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 6 citations
- DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable ArraysJiayi Wang, Ang Da Lu, Zhichen Zeng, Ang LiISCA 2026
