sCROOGe: Circuit-level Design and Optimization Framework for RISC-V Out-of-Order GPUs
Maria Zerva, Panagiotis-Eleftherios Eleftherakis, Alexis Maras, Konstantinos Iliakis, Alexandros Moiras, Sotirios Xydis
Abstract
Graphics Processing Units (GPUs) have evolved into the dominant hardware accelerators for general-purpose computing, yet many workloads underutilize the available resources due to inadequate Thread-Level Parallelism (TLP). To address this, techniques leveraging Instruction-Level Parallelism (ILP), such as dynamic instruction reordering, have been proposed. However, existing solutions rely on software simulators lacking RTL validation and abstracting away critical micro-architectural details, limiting accuracy in performance, power, and area modeling. In this work, we present the first synthesizable RTL assessment and optimization framework of both frontend- and backend-based Out-of-Order (OoO) execution schemes within the open-source RISC-V Vortex GPGPU framework. Both schemes are directly implemented in RTL by extending the pipeline with light-weight scheduling logic and register renaming. Our approach manages to capture critical micro-architectural details absent from prior simulation-only studies and reveals key insights into the performance scalability and implementation cost of OoO execution paths in GPUs. We leverage this flexibility to explore different reordering configurations, isolate the impact of key components, and optimize hardware structures for balanced performance, power, area and timing. We evaluate their performance across diverse workloads, and perform design space exploration by tuning parameters such as warp and thread counts. Furthermore, we quantify the power and area trade-offs via ASIC synthesis flow, demonstrating a 14.4% performance gain compared to iso-area in-order GPU cores and 27.9% improved Energy-Delay Product (EDP). This work demonstrates the practical applicability of light-weight OoO schemes for enhanced GPU throughput and energy efficiency, establishing a foundation for future ILP-aware designs validated through real hardware modeling.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e9023793-254c-4d7d-88ca-065abdb9da2dRelated papers
- Vortex: Extending the RISC-V ISA for GPGPU and 3D-GraphicsBlaise Tine, Krishna Praveen Yalamarthy, Fares Elsabbagh, Hyesoon KimMICRO 2021 · 61 citations
- SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsShinnung Jeong, Liam Paul Cooper, Ju Min Lee, Heelim Choi et al.HPCA 2025 · 2 citations
- Fuzzing Open-Source GPU Hardware with SIMT Program GenerationZibo Gao, Jie Wang, Qihang Zhou, Lixiao Shan et al.USENIX Security 2026
- Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data AnalyticsYichao Yuan, Advait Iyer, Lin Ma, Nishil TalatiVLDB 2025 · 11 citations
- GhOST: a GPU Out-of-Order Scheduling Technique for Stall ReductionIshita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu et al.ISCA 2024 · 9 citations
