Ultra-Elastic CGRAs for Irregular Loop Specialization
Christopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan, Christopher Batten
Abstract
Reconfigurable accelerator fabrics, including coarse-grain reconfigurable arrays (CGRAs), have experienced a resurgence in interest because they allow fast-paced software algorithm development to continue evolving post-fabrication. CGRAs traditionally target regular workloads with data-level parallelism (e.g., neural networks, image processing), but once integrated into an SoC they remain idle and unused for irregular workloads. An emerging trend towards repurposing these idle resources raises important questions for how to efficiently map and execute general-purpose loops which may have irregular memory accesses, irregular control flow, and inter-iteration loop dependencies. Recent work has increasingly leveraged elasticity in CGRAs to mitigate the first two challenges, but elasticity alone does not address inter-iteration loop dependencies which can easily bottleneck overall performance. In this paper, we address all three challenges for irregular loop specialization and propose ultra-elastic CGRAs (UE-CGRAs), a novel elastic CGRA that accelerates true-dependency bottlenecks and saves energy in irregular loops by overcoming traditional VLSI challenges. UE-CGRAs allow configurable fine-grain dynamic voltage and frequency scaling (DVFS) for each of potentially hundreds of tiny processing elements (PEs) in the CGRA, enabling chains of connected PEs to “rest” at lower voltages and frequencies to save energy, while other chains of connected PEs can “sprint” at higher voltages and frequencies to accelerate through true-dependency bottlenecks. UE-CGRAs rely on a novel ratiochronous clocking scheme carefully overlaid on the inter-PE elastic interconnect to enable low-latency crossings while remaining fully verifiable with commercial static timing analysis tools. We present the UE-CGRA analytical model, compiler, architectural template, and VLSI circuitry, and we demonstrate how UE-CGRAs can specialize for irregular loops and improve performance () or energy efficiency with reasonable area overhead compared to traditional inelastic and elastic CGRAs, while also improving performance () or energy efficiency (up to ) compared to a RISC-V core.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e728cf41-553e-44d7-84a0-0488d520f29eCited by top-tier papers12
- LISA: Graph Neural Network based Portable Mapping on Spatial AcceleratorsZhaoying Li, Dan Wu, Dhananjaya Wijerathne, Tulika MitraHPCA 2022 · 43 citations
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
- SCALO: An Accelerator-Rich Distributed System for Scalable Brain-Computer InterfacingKarthik Sriram, Raghavendra Pradyumna Pothukuchi, Michal Gerasimiuk, Muhammed Ugur et al.ISCA 2023 · 18 citations
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang et al.ASPLOS 2025 · 17 citations
- Pipestitch: An energy-minimal dataflow architecture with lightweight threadsNathan Serafin, Souradip Ghosh, Harsh Desai, Nathan Beckmann et al.MICRO 2023 · 16 citations
Related papers
- ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationCheng Tan, Miaomiao Jiang, Deepak Patil, Yanghui Ou et al.MICRO 2024 · 10 citations
- TAEM: Fast Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAMingyang Kou, Jiangyuan Gu, Shaojun Wei, Hailong Yao et al.DAC 2020 · 18 citations
- GEML: GNN-based efficient mapping method for large loop applications on CGRAMingyang Kou, Jun Zeng, Boxiao Han, Fei Xu et al.DAC 2022 · 17 citations
- Optimizing Data Reuse for CGRA Mapping Using Polyhedral-based Loop TransformationsLiao Huang, Dajiang LiuDAC 2023 · 4 citations
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia et al.ISCA 2021 · 84 citations
