BOW: Breathing Operand Windows to Exploit Bypassing in GPUs
Hodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong, Nael B. Abu-Ghazaleh
Abstract
The Register File (RF) is a critical structure in Graphics Processing Units (GPUs) responsible for a large portion of the area and power. To simplify the architecture of the RF, it is organized in a multi-bank configuration with a single port for each bank. Not surprisingly, the frequent accesses to the register file during kernel execution incur a sizeable overhead in GPU power consumption, and introduce delays as accesses are serialized when port conflicts occur. In this paper, we observe that there is a high degree of temporal locality in accesses to the registers: within short instruction windows, the same registers are often accessed repeatedly. We characterize the opportunities to reduce register accesses as a function of the size of the instruction window considered, and establish that there are many recurring reads and updates of the same register operands in most GPU computations. To exploit this opportunity, we propose Breathing Operand Windows (BOW), an enhanced GPU pipeline and operand collector organization that supports bypassing register file accesses and instead passes values directly between instructions within the same window. Our baseline design can only bypass register reads; we introduce an improved design capable of also bypassing unnecessary write operations to the RF. We introduce compiler optimizations to help guide the write-back destination of operands depending on whether they will be reused to further reduce the write traffic. To reduce the storage overhead, we analyze the occupancy of the bypass buffers and discover that we can significantly down size them without losing performance. BOW along with optimizations reduces dynamic energy consumption of the register file by 55% and increases the performance by 11%, with a modest overhead of 12KB increase in the size of the operand collectors (4% of the register file size).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd3805e5-ab9a-4cb0-8b46-42c17182f19bCited by top-tier papers2
- GhOST: a GPU Out-of-Order Scheduling Technique for Stall ReductionIshita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu et al.ISCA 2024 · 9 citations
- Dissecting and Modeling the Architecture of Modern GPU CoresRodrigo Huerta, Mojtaba Abaie Shoushtary, José-Lorenzo Cruz, Antonio GonzálezMICRO 2025 · 8 citations
Builds on1
Related papers
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 2 citations
- Exploiting Zero Data to Reduce Register File and Execution Unit Dynamic Power Consumption in GPGPUsAhmad M. Radaideh, Paul V. GratzDAC 2020 · 4 citations
- Memento: An Adaptive, Compiler-Assisted Register File Cache for GPUsMojtaba Abaie Shoushtary, José-María Arnau, Jordi Tubella Murgadas, Antonio GonzálezISCA 2024 · 4 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim et al.MICRO 2020 · 27 citations
