Lune

ISCA2026Top-tier venue

DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays

Jiayi Wang, Ang Da Lu, Zhichen Zeng, Ang Li

2026Year

Abstract

While general-purpose graphics processing units (GPGPUs) dominate massively parallel computing through the single-instruction, multiple-thread (SIMT) programming model, their underlying single-instruction, multiple-data (SIMD) execution incurs substantial energy overhead from frequent register file (RF) accesses and complex control logic. We present DICE, a novel architecture that addresses these inefficiencies by replacing the SIMD backend with minimaloverhead, statically scheduled coarse-grained reconfigurable arrays (CGRAs) while preserving the well-known SIMT programming model. Unlike SIMD units that execute warps of threads in lockstep, DICE dispatches active threads in a pipelined manner onto the CGRA fabric, where data flows directly between processing elements (PEs), reducing RF accesses for intermediate values. Moreover, DICE organizes threads into groups larger than warps, reducing overhead by eliminating the warp-level control logic and further amortizing remaining control overhead across hundreds of threads. To handle operations with runtime dynamism, such as variable-latency memory loads and data-dependent control flow, while preserving static scheduling, DICE compiles programs into p-graphs by partitioning dynamic dependence edges across separate CGRA configurations. DICE further introduces several key optimizations: double-buffered configuration memory to hide reconfiguration latency, compile-time pp-graph unrolling to enhance resource utilization, and a temporal memory coalescing unit (TMCU) to merge memory requests from consecutive, pipelined threads. Evaluations on Rodinia benchmarks in Accelsim demonstrate that DICE reduces register file accesses by 68% on average. With equivalent computation and memory resources, DICE's CGRA Processors (CPs) achieve a geometric mean of 1. 7 7 - 1. 9 0×\text{1. 7 7 - 1. 9 0} \times dynamic energy efficiency and 4 2. 0 % - 4 5. 9 %\text{4 2. 0 \% - 4 5. 9 \%} average power reduction compared to the modeled NVIDIA Turing Streaming Multiprocessors (SMs), while the full DICE system achieves performance comparable to the modeled Turing GPU baselines. DICE demonstrates that spatial pipeline execution can deliver substantial energy savings without sacrificing performance.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 03544817-2ff6-4cca-ba5b-2821950132cb

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines