TaskStream: accelerating task-parallel workloads by recovering program structure
Vidushi Dadu, Tony Nowatzki
Abstract
Reconfigurable accelerators, like CGRAs and dataflow architectures, have come to prominence for addressing data-processing problems. However, they are largely limited to workloads with regular parallelism, precluding their applicability to prevalent task-parallel workloads. Reconfigurable architectures and task parallelism seem to be at odds, as the former requires repetitive and simple program structure, and the latter breaks program structure to create small, individually scheduled program units.
Our insight is that if tasks and their potential for communication structure are first-class primitives in the hardware, it is possible to recover program structure with extremely low overhead. We propose a task execution model for accelerators called TaskStream, which annotates task dependences with information sufficient to recover inter-task structure. TaskStream enables work-aware load balancing, recovery of pipelined inter-task dependences, and recovery of inter-task read sharing through multicasting.
We apply TaskStream to a reconfigurable dataflow architecture, creating a seamless hierarchical dataflow model for task-parallel workloads. We compare our accelerator, Delta, with an equivalent static-parallel design. Overall, we find that our execution model can improve performance by 2.2× with only 3.6% area overhead, while alleviating the programming burden of managing task distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af8f8d8a-d217-4a98-8028-2912fc2a6cffCited by top-tier papers9
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh et al.MICRO 2022 · 32 citations
- ABNDP: Co-optimizing Data Access and Load Balance in Near-Data ProcessingBoyu Tian, Qihang Chen, Mingyu GaoASPLOS 2023 · 31 citations
- Profiling Hyperscale Big Data ProcessingAbraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu et al.ISCA 2023 · 30 citations
- Spatula: A Hardware Accelerator for Sparse Matrix FactorizationAxel Feldmann, Daniel SánchezMICRO 2023 · 9 citations
- Enhancing CGRA Efficiency Through Aligned Compute and Communication ProvisioningZhaoying Li, Pranav Dangi, Chenyang Yin, Thilini Kaushalya Bandara et al.ASPLOS 2025 · 8 citations
Builds on14
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationGuowei Zhang, Nithya Attaluri, Joel S. Emer, Daniel SánchezASPLOS 2021 · 158 citations
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
- Ultra-Elastic CGRAs for Irregular Loop SpecializationChristopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan et al.HPCA 2021 · 68 citations
Related papers
- Streaming Task Graph Scheduling for Dataflow ArchitecturesTiziano De Matteis, Lukas Gianinazzi, Johannes de Fine Licht, Torsten HoeflerHPDC 2023 · 3 citations
- NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable ArchitecturesShangkun Li, Jinming Ge, Diyuan Tao, Zeyu Li et al.PLDI 2026
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 4 citations
- PANORAMA: divide-and-conquer approach for mapping complex loop kernels on CGRADhananjaya Wijerathne, Zhaoying Li, Thilini Kaushalya Bandara, Tulika MitraDAC 2022 · 18 citations
- TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAsNeha Prakriya, Yuze Chi, Suhail Basalama, Linghao Song et al.ASPLOS 2024 · 6 citations
