ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAs
Yixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li, Antonino Tumeo, Nirav Dave, Tong Geng
摘要
Coarse-Grained Reconfigurable Arrays (CGRAs) can achieve higher energy-efficiency than general-purpose processors and accelerators or fine-grained reconfigurable devices, while maintaining adaptability to different computational patterns. CGRAs have shown some success as a platform to accelerate machine learning (ML) thanks to their flexibility, which allows them to support new models not considered by fixed accelerators. However, current solutions for CGRAs employ low level instruction-based compiler approaches and lack specialized compilation infrastructures from high-level ML frameworks that could leverage semantic information from the models, limiting the ability to efficiently map them on the reconfigurable substrate. This paper proposes ML-CGRA, an integrated compilation framework based on the MLIR infrastructure that enables efficient ML acceleration on CGRAs. ML-CGRA provides an end-to-end solution for mapping ML models on CGRAs that outperforms conventional approaches by 3.15× and 6.02 × on 4×4 and 8×8 CGRAs, respectively. The framework is open-source and available from https://github.com/tancheng/mlir-cgra.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang 等ASPLOS 2025 · 被引用 17 次
- Enhancing CGRA Efficiency Through Aligned Compute and Communication ProvisioningZhaoying Li, Pranav Dangi, Chenyang Yin, Thilini Kaushalya Bandara 等ASPLOS 2025 · 被引用 8 次
- TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow AcceleratorsWei Li, Zhenyu Bai, Heru Wang, Pranav Dangi 等OSDI 2026 · 被引用 2 次
- NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable ArchitecturesShangkun Li, Jinming Ge, Diyuan Tao, Zeyu Li 等PLDI 2026
- GenZA: A General and Efficient Accelerator for Diverse Zero-Knowledge Proof ProtocolsCheng Wang, Jiangbin Dong, Mingyu GaoISCA 2026
相关 Paper
- E2EMap: End-to-End Reinforcement Learning for CGRA Compilation via Reverse MappingDajiang Liu, Yuxin Xia, Jiaxing Shang, Jiang Zhong 等HPCA 2024 · 被引用 14 次
- APEX: A Framework for Automated Processing Element Design Space Exploration using Frequent Subgraph AnalysisJackson Melchert, Kathleen Feng, Caleb Donovick, Ross Daly 等ASPLOS 2023 · 被引用 12 次
- Ultra-Fast CGRA Scheduling to Enable Run Time, Programmable CGRAsJinho Lee, Trevor E. CarlsonDAC 2021 · 被引用 16 次
- FHE-CGRA: Enable Efficient Acceleration of Fully Homomorphic Encryption on CGRAsMiaomiao Jiang, Yilan Zhu, Honghui You, Cheng Tan 等DAC 2024 · 被引用 5 次
- MapZero: Mapping for Coarse-grained Reconfigurable Architectures with Reinforcement Learning and Monte-Carlo Tree SearchXiangyu Kong, Yi Huang, Jianfeng Zhu, Xingchen Man 等ISCA 2023 · 被引用 30 次
