NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures
Shangkun Li, Jinming Ge, Diyuan Tao, Zeyu Li, Jiawei Liang, Linfeng Du, Jiang Xu, Wei Zhang, Cheng Tan
Abstract
Coarse-Grained Reconfigurable Architectures (CGRAs) are a promising and versatile accelerator platform, offering a balance between the performance and efficiency of specialized accelerators and the software programmability. However, their full potential is severely hindered by control flow in accelerated kernels, as the control flow (e.g., loops, branches) is fundamentally incompatible with the parallel, data-driven CGRA fabric. Prior strategies to resolve this mismatch in CGRA kernel acceleration are either inefficient, sacrificing performance for generality, or lack generality due to the difficulty of adapting them across different execution models. Thus, a general and unified solution for efficient CGRA kernel acceleration remains elusive.
This paper introduces NEURA, a unified and retargetable compilation framework that systematically resolves the control-dataflow mismatch in CGRAs. NEURA's core innovation is a novel, pure dataflow intermediate representation (IR) built on a predicated type system. In this IR, control contexts are embedded as a predicate within each data, making control an intrinsic property of data. This mechanism enables NEURA to systematically flatten complex control flow into a single unified dataflow graph. This unified representation decouples kernel representation from hardware, empowering NEURA to retarget diverse CGRAs with different execution models and microarchitectural features. When targeted to a high-performance spatio-temporal CGRA, NEURA delivers a 2.20× speedup on kernel benchmarks and up to 2.71× geometric mean speedup on real-world applications over state-of-the-art (SOTA) high-performance baselines. It also provides a competitive solution against the SOTA low-power CGRA when retargeted to a spatial-only CGRA. NEURA is open-source and available at https://github.com/coredac/neura.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a44956dc-e1eb-4222-baf9-d075c158f49bBuilds on19
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia et al.ISCA 2021 · 84 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
- Ultra-Elastic CGRAs for Irregular Loop SpecializationChristopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan et al.HPCA 2021 · 68 citations
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry et al.MICRO 2022 · 66 citations
Related papers
- ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAsYixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li et al.DAC 2023 · 39 citations
- PANORAMA: divide-and-conquer approach for mapping complex loop kernels on CGRADhananjaya Wijerathne, Zhaoying Li, Thilini Kaushalya Bandara, Tulika MitraDAC 2022 · 18 citations
- LISA: Graph Neural Network based Portable Mapping on Spatial AcceleratorsZhaoying Li, Dan Wu, Dhananjaya Wijerathne, Tulika MitraHPCA 2022 · 43 citations
- Adora Compiler: End-to-End Optimization for High-Efficiency Dataflow Acceleration and Task Pipelining on CGRAsJiahang Lou, Qilong Zhu, Yuan Dai, Zewei Zhong et al.DAC 2025 · 2 citations
- TAEM: Fast Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAMingyang Kou, Jiangyuan Gu, Shaojun Wei, Hailong Yao et al.DAC 2020 · 18 citations
