Pipestitch: An energy-minimal dataflow architecture with lightweight threads
Nathan Serafin, Souradip Ghosh, Harsh Desai, Nathan Beckmann, Brandon Lucia
摘要
Computing at the extreme edge allows systems with high-resolution sensors to be pushed well outside the reach of traditional communication and power delivery, requiring high-performance, high-energy-efficiency architectures to run complex ML, DSP, image processing, etc. Recent work has demonstrated the suitability of CGRAs for energy-minimal computation, but has focused strictly on energy optimization, neglecting performance. Pipestitch is an energy-minimal CGRA architecture that adds lightweight hardware threads to ordered dataflow, exploiting abundant, untapped parallelism in the complex workloads needed to meet the demands of emerging sensing applications. Pipestitch introduces a programming model, control-flow operator, and synchronization network to allow lightweight hardware threads to pipeline on the CGRA fabric. Across 5 important sparse workloads, Pipestitch achieves a 3.49 × increase in performance over RipTide, the state-of-the-art, at a cost of a 1.10 × increase in area and a 1.05 × increase in energy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang 等ASPLOS 2025 · 被引用 17 次
- Enhancing CGRA Efficiency Through Aligned Compute and Communication ProvisioningZhaoying Li, Pranav Dangi, Chenyang Yin, Thilini Kaushalya Bandara 等ASPLOS 2025 · 被引用 8 次
- FlowCert: Translation Validation for Asynchronous Dataflow via Dynamic Fractional PermissionsZhengyao Lin, Joshua Gancher, Bryan ParnoOOPSLA 2024 · 被引用 4 次
- Ripple: Asynchronous Programming for Spatial Dataflow ArchitecturesSouradip Ghosh, Yufei Shi, Brandon Lucia, Nathan BeckmannPLDI 2025 · 被引用 4 次
- The TYR Dataflow Architecture: Improving Locality by Taming ParallelismNikhil Agarwal, Mitchell Fream, Souradip Ghosh, Brian C. Schwedock 等MICRO 2024 · 被引用 1 次
它引用的顶会 Paper15
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang 等NeurIPS 2022 · 被引用 345 次
- Orbital Edge Computing: Nanosatellite Constellations as a New Class of Computer SystemBradley Denby, Brandon LuciaASPLOS 2020 · 被引用 272 次
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi 等MICRO 2020 · 被引用 223 次
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 被引用 158 次
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
相关 Paper
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry 等MICRO 2022 · 被引用 66 次
- Adora Compiler: End-to-End Optimization for High-Efficiency Dataflow Acceleration and Task Pipelining on CGRAsJiahang Lou, Qilong Zhu, Yuan Dai, Zewei Zhong 等DAC 2025 · 被引用 2 次
- DRIPS: Dynamic Rebalancing of Pipelined Streaming Applications on CGRAsCheng Tan, Nicolas Bohm Agostini, Tong Geng, Chenhao Xie 等HPCA 2022 · 被引用 28 次
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 被引用 60 次
- DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable ArraysJiayi Wang, Ang Da Lu, Zhichen Zeng, Ang LiISCA 2026
