APEX: A Framework for Automated Processing Element Design Space Exploration using Frequent Subgraph Analysis
Jackson Melchert, Kathleen Feng, Caleb Donovick, Ross Daly, Ritvik Sharma, Clark W. Barrett, Mark A. Horowitz, Pat Hanrahan, Priyanka Raina
Abstract
The architecture of a coarse-grained reconfigurable array (CGRA) processing element (PE) has a significant effect on the performance and energy-efficiency of an application running on the CGRA. This paper presents APEX, an automated approach for generating specialized PE architectures for an application or an application domain. APEX first analyzes application domain benchmarks using frequent subgraph mining to extract commonly occurring computational subgraphs. APEX then generates specialized PEs by merging subgraphs using a datapath graph merging algorithm. The merged datapath graphs are translated into a PE specification from which we automatically generate the PE hardware description in Verilog along with a compiler that maps applications to the PE. The PE hardware and compiler are inserted into a flexible CGRA generation and compilation toolchain that allows for agile evaluation of CGRAs. We evaluate APEX for two domains, machine learning and image processing. For image processing applications, our automatically generated CGRAs with specialized PEs achieve from 5% to 30% less area and from 22% to 46% less energy compared to a general-purpose CGRA. For machine learning applications, our automatically generated CGRAs consume 16% to 59% less energy and 22% to 39% less area than a general-purpose CGRA. This work paves the way for creation of application domain-driven design-space exploration frameworks that automatically generate efficient programmable accelerators, with a much lower design effort for both hardware and compiler generation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 27a269a5-bfaf-470b-bf47-2d26c653a054Cited by top-tier papers3
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang et al.ASPLOS 2025 · 17 citations
- ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationCheng Tan, Miaomiao Jiang, Deepak Patil, Yanghui Ou et al.MICRO 2024 · 10 citations
- Enhancing CGRA Efficiency Through Aligned Compute and Communication ProvisioningZhaoying Li, Pranav Dangi, Chenyang Yin, Thilini Kaushalya Bandara et al.ASPLOS 2025 · 8 citations
Related papers
- ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAsYixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li et al.DAC 2023 · 39 citations
- E2EMap: End-to-End Reinforcement Learning for CGRA Compilation via Reverse MappingDajiang Liu, Yuxin Xia, Jiaxing Shang, Jiang Zhong et al.HPCA 2024 · 14 citations
- Mixed-granularity parallel coarse-grained reconfigurable architectureJinyi Deng, Linyun Zhang, Lei Wang, Jiawei Liu et al.DAC 2022 · 5 citations
- MapZero: Mapping for Coarse-grained Reconfigurable Architectures with Reinforcement Learning and Monte-Carlo Tree SearchXiangyu Kong, Yi Huang, Jianfeng Zhu, Xingchen Man et al.ISCA 2023 · 30 citations
- Ultra-Fast CGRA Scheduling to Enable Run Time, Programmable CGRAsJinho Lee, Trevor E. CarlsonDAC 2021 · 16 citations
