DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
Jiayi Wang, Ang Da Lu, Zhichen Zeng, Ang Li
Abstract
While general-purpose graphics processing units (GPGPUs) dominate massively parallel computing through the single-instruction, multiple-thread (SIMT) programming model, their underlying single-instruction, multiple-data (SIMD) execution incurs substantial energy overhead from frequent register file (RF) accesses and complex control logic. We present DICE, a novel architecture that addresses these inefficiencies by replacing the SIMD backend with minimaloverhead, statically scheduled coarse-grained reconfigurable arrays (CGRAs) while preserving the well-known SIMT programming model. Unlike SIMD units that execute warps of threads in lockstep, DICE dispatches active threads in a pipelined manner onto the CGRA fabric, where data flows directly between processing elements (PEs), reducing RF accesses for intermediate values. Moreover, DICE organizes threads into groups larger than warps, reducing overhead by eliminating the warp-level control logic and further amortizing remaining control overhead across hundreds of threads. To handle operations with runtime dynamism, such as variable-latency memory loads and data-dependent control flow, while preserving static scheduling, DICE compiles programs into p-graphs by partitioning dynamic dependence edges across separate CGRA configurations. DICE further introduces several key optimizations: double-buffered configuration memory to hide reconfiguration latency, compile-time -graph unrolling to enhance resource utilization, and a temporal memory coalescing unit (TMCU) to merge memory requests from consecutive, pipelined threads. Evaluations on Rodinia benchmarks in Accelsim demonstrate that DICE reduces register file accesses by 68% on average. With equivalent computation and memory resources, DICE's CGRA Processors (CPs) achieve a geometric mean of dynamic energy efficiency and average power reduction compared to the modeled NVIDIA Turing Streaming Multiprocessors (SMs), while the full DICE system achieves performance comparable to the modeled Turing GPU baselines. DICE demonstrates that spatial pipeline execution can deliver substantial energy savings without sacrificing performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03544817-2ff6-4cca-ba5b-2821950132cbBuilds on5
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan et al.MICRO 2021 · 134 citations
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia et al.ISCA 2021 · 84 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
- Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated SupercomputingOscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding et al.SC 2025 · 5 citations
Related papers
- ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationCheng Tan, Miaomiao Jiang, Deepak Patil, Yanghui Ou et al.MICRO 2024 · 10 citations
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 60 citations
- Dimensionality-Aware Redundant SIMT Instruction EliminationTsung Tai Yeh, Roland N. Green, Timothy G. RogersASPLOS 2020 · 12 citations
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry et al.MICRO 2022 · 66 citations
- DARIC: A Data Reuse-Friendly CGRA for Parallel Data Access via Elastic FIFOsDajiang Liu, Di Mou, Rong Zhu, Yan Zhuang et al.DAC 2023 · 7 citations
