Fifer: Practical Acceleration of Irregular Applications on Reconfigurable Architectures
Quan M. Nguyen, Daniel Sánchez
摘要
Coarse-grain reconfigurable arrays (CGRAs) can achieve much higher performance and efficiency than general-purpose cores, approaching the performance of a specialized design while retaining programmability. Unfortunately, CGRAs have so far only been effective on applications with regular compute patterns. However, many important workloads like graph analytics, sparse linear algebra, and databases, are irregular applications with unpredictable access patterns and control flow. Since CGRAs map computation statically to a spatial fabric of functional units, irregular memory accesses and control flow cause frequent stalls and load imbalance.
We present Fifer, an architecture and compilation technique that makes irregular applications efficient on CGRAs. Fifer first decouples irregular applications into a feed-forward network of pipeline stages. Each resulting stage is regular and can efficiently use the CGRA fabric. However, irregularity causes stages to have widely varying loads, resulting in high load imbalance if they execute spatially in a conventional CGRA. Fifer solves this by introducing dynamic temporal pipelining: it time-multiplexes multiple stages onto the same CGRA, and dynamically schedules stages to avoid load imbalance. Fifer makes time-multiplexing fast and cheap to quickly respond to load imbalance while retaining the efficiency and simplicity of a CGRA design. We show that Fifer improves performance by gmean 2.8× (and up to 5.5×) over a conventional CGRA architecture (and by gmean 17× over an out-of-order multicore) on a variety of challenging irregular applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- The Sparse Abstract MachineOlivia Hsu, Maxwell Strange, Ritvik Sharma, Jaeyeon Won 等ASPLOS 2023 · 被引用 37 次
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh 等MICRO 2022 · 被引用 32 次
- Dalorex: A Data-Local Program Execution and Architecture for Memory-bound ApplicationsMarcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret MartonosiHPCA 2023 · 被引用 23 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
- Pipestitch: An energy-minimal dataflow architecture with lightweight threadsNathan Serafin, Souradip Ghosh, Harsh Desai, Nathan Beckmann 等MICRO 2023 · 被引用 16 次
它引用的顶会 Paper4
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 被引用 28 次
- Gorgon: Accelerating Machine Learning from Relational DataMatthew Vilim, Alexander Rucker, Yaqi Zhang, Sophia Liu 等ISCA 2020 · 被引用 25 次
相关 Paper
- ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAsYixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li 等DAC 2023 · 被引用 39 次
- TAEM: Fast Transfer-Aware Effective Loop Mapping for Heterogeneous Resources on CGRAMingyang Kou, Jiangyuan Gu, Shaojun Wei, Hailong Yao 等DAC 2020 · 被引用 18 次
- Rewire: Advancing CGRA Mapping Through a Consolidated Routing ParadigmZhaoying Li, Dan Wu, Dhananjaya Wijerathne, Dan Chen 等DAC 2025
- Ultra-Fast CGRA Scheduling to Enable Run Time, Programmable CGRAsJinho Lee, Trevor E. CarlsonDAC 2021 · 被引用 16 次
- DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAsCheng Tan, Nicolas Bohm Agostini, Tong Geng, Chenhao Xie 等HPCA 2022 · 被引用 28 次
