Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand Compaction
Eunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung Kim
摘要
The GPU has been successfully used for diverse emerging compute-intensive applications, including imaging, computer vision, and more recently, deep learning, to name a few. To offer high performance for such applications, it is provisioned with massive Register Files (RFs) to exploit high Thread-Level Parallelism (TLP). As RFs are designed to store thousands of contexts and provide very high bandwidth for uninterrupted supply of operands to hundreds of compute units for high TLP, they have become one of the most power-hungry components in the GPU. Meanwhile, faced with the end of Dennard scaling, it is more important than ever to innovate the microarchitecture of (power-constrained) GPUs to continuously improve the performance of future compute-intensive applications.In this work, we propose a GPU microarchitecture, WarpedCompaction, designed to utilize given RFs more efficiently, instead of relying on larger-capacity and/or higher-bandwidth RFs to further improve GPU performance. Specifically, first, we reverse-engineer the latest GPU’s RF organization through microbenchmarking. This uncovers that each sub-core within a streaming multiprocessor contains only two dual-ported RF banks, and accesses to these banks are arbitrated solely based on register IDs. We also reveal that, despite the modest configurations, RF banks are largely underutilized, staying inactive for 33.5% of the time. Second, we observe that previously proposed RF optimization techniques, data forwarding and dead register elimination, cannot address this underutilization problem. This is mainly due to the (R1) insufficient RF access requests from a limited number of Operand Collector Units (OCUs) and (R2) inefficient operand distribution by the conventional RF bank arbitration. Third, we present two architectural solutions to tackle the observed inefficiency: (S1) OCU sharing and (S2) skewed arbitrator. Building on enhanced OCU early allocation for partially ready instructions and operand forwarding to OCUs, OCU sharing allows two different warp instructions to share a single OCU, and skewed arbitrator evenly distributes register accesses across RF banks, maximizing the utilization of given RF bandwidth. The synergistic integration of these techniques, forming Warped-Compaction, results in higher performance and better energy efficiency of RF and OCU compared to baseline high-end GPU.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- BOW: Breathing Operand Windows to Exploit Bypassing in GPUsHodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong 等MICRO 2020 · 被引用 19 次
- Mitigating GPU Core Partitioning Performance EffectsAaron Barnes, Fangjia Shen, Timothy G. RogersHPCA 2023 · 被引用 18 次
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 被引用 13 次
- Memento: An Adaptive, Compiler-Assisted Register File Cache for GPUsMojtaba Abaie Shoushtary, José-María Arnau, Jordi Tubella Murgadas, Antonio GonzálezISCA 2024 · 被引用 4 次
- ACRS: Adjacent Computation Resource Sharing among Partitioned GPU Sub-CoresPenghao Song, Chongxi Wang, Chenji Han, Haoyu Zhao 等DAC 2025
