Dimensionality-Aware Redundant SIMT Instruction Elimination
Tsung Tai Yeh, Roland N. Green, Timothy G. Rogers
Abstract
In massively multithreaded architectures, redundantly executing the same instruction with the same operands in different threads is a significant source of inefficiency. This paper introduces Dimensionality-Aware Redundant SIMT Instruction Elimination (DARSIE), a non-speculative instruction skipping mechanism to reduce redundant operations in GPUs. DARSIE uses static markings from the compiler and information obtained at kernel launch time to skip redundant instructions before they are fetched, keeping them out of the pipeline. DARSIE exploits a new observation that there is significant redundancy across warp instructions in multi-dimensional threadblocks.
For minimal area cost, DARSIE eliminates conditionally redundant instructions without any programmer intervention. On increasingly important 2D GPU applications, DARSIE reduces the number of instructions fetched and executed by 23% over contemporary GPUs. Not fetching these instructions results in a geometric mean of 30% performance improvement, while decreasing the energy consumed by 25%.
• Computer systems organization → Single instruction, multiple data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6095c86-0782-4eac-ae55-00d5939b59b4Cited by top-tier papers3
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim et al.MICRO 2020 · 27 citations
- GVProf: a value profiler for GPU-based clustersKeren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng et al.SC 2020 · 24 citations
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 3 citations
Related papers
- R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUsDongho Ha, Yunho Oh, Won Woo RoISCA 2023 · 9 citations
- DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable ArraysJiayi Wang, Ang Da Lu, Zhichen Zeng, Ang LiISCA 2026
- RedSan: A Redundant Memory Instruction Sanitizer for GPU ProgramsYanbo Zhao, Yueming Hao, Zecheng Li, Shuyin Jiao et al.SC 2025 · 1 citation
- VISTA: Optimizing GPU Scheduling through Versatile Locality-Aware Data SharingHajar Falahati, Negin Mahani, Adrián Cristal, Osman S. UnsalDAC 2025 · 1 citation
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
