DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
Jiayi Wang, Ang Da Lu, Zhichen Zeng, Ang Li
摘要
While general-purpose graphics processing units (GPGPUs) dominate massively parallel computing through the single-instruction, multiple-thread (SIMT) programming model, their underlying single-instruction, multiple-data (SIMD) execution incurs substantial energy overhead from frequent register file (RF) accesses and complex control logic. We present DICE, a novel architecture that addresses these inefficiencies by replacing the SIMD backend with minimaloverhead, statically scheduled coarse-grained reconfigurable arrays (CGRAs) while preserving the well-known SIMT programming model. Unlike SIMD units that execute warps of threads in lockstep, DICE dispatches active threads in a pipelined manner onto the CGRA fabric, where data flows directly between processing elements (PEs), reducing RF accesses for intermediate values. Moreover, DICE organizes threads into groups larger than warps, reducing overhead by eliminating the warp-level control logic and further amortizing remaining control overhead across hundreds of threads. To handle operations with runtime dynamism, such as variable-latency memory loads and data-dependent control flow, while preserving static scheduling, DICE compiles programs into p-graphs by partitioning dynamic dependence edges across separate CGRA configurations. DICE further introduces several key optimizations: double-buffered configuration memory to hide reconfiguration latency, compile-time -graph unrolling to enhance resource utilization, and a temporal memory coalescing unit (TMCU) to merge memory requests from consecutive, pipelined threads. Evaluations on Rodinia benchmarks in Accelsim demonstrate that DICE reduces register file accesses by 68% on average. With equivalent computation and memory resources, DICE's CGRA Processors (CPs) achieve a geometric mean of dynamic energy efficiency and average power reduction compared to the modeled NVIDIA Turing Streaming Multiprocessors (SMs), while the full DICE system achieves performance comparable to the modeled Turing GPU baselines. DICE demonstrates that spatial pipeline execution can deliver substantial energy savings without sacrificing performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan 等MICRO 2021 · 被引用 134 次
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia 等ISCA 2021 · 被引用 84 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
- Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated SupercomputingOscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 等SC 2025 · 被引用 5 次
相关 Paper
- ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationCheng Tan, Miaomiao Jiang, Deepak Patil, Yanghui Ou 等MICRO 2024 · 被引用 10 次
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 被引用 60 次
- Dimensionality-Aware Redundant SIMT Instruction EliminationTsung Tai Yeh, Roland N. Green, Timothy G. RogersASPLOS 2020 · 被引用 12 次
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry 等MICRO 2022 · 被引用 66 次
- DARIC: A Data Reuse-Friendly CGRA for Parallel Data Access via Elastic FIFOsDajiang Liu, Di Mou, Rong Zhu, Yan Zhuang 等DAC 2023 · 被引用 7 次
