A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC Workloads
Nicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda
摘要
Dataflow accelerators can provide energy efficient and high-performance alternatives to current popular architectures. However, little work has been done to enable accelerator-initiated, scalable collective communication for these architectures. We develop a High Level Synthesis (HLS) interface to bridge this gap through software-hardware co-design. Given the tendency of dataflow applications to use reads and writes to streams to express data transfer, we develop a streaming interface implementing fine-grained transfers to the host processor. Data can then be communicated through MPI and transferred to the receiving accelerator. As a result, the interface uses few hardware resources for communication. As a case study, we enhance the HPL benchmark from HPCC_FPGA with our contributions. We evaluate our final design on up to 16 FPGAs, achieving up to 18% improvement in application throughput and 36% reduced latency in kernel execution. Additionally, we design a stencil benchmark that showcases superlinear speedup in a strong scaling scenario.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CODO: An Automated Compiler for Comprehensive Dataflow OptimizationWeichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang 等ISCA 2026
- SOFF: An OpenCL High-Level Synthesis Framework for FPGAsGangwon Jo, Heehoon Kim, Jeesoo Lee, Jaejin LeeISCA 2020 · 被引用 20 次
- fBLAS: streaming linear algebra on FPGATiziano De Matteis, Johannes de Fine Licht, Torsten HoeflerSC 2020 · 被引用 27 次
- HIDA: A Hierarchical Dataflow Compiler for High-Level SynthesisHanchen Ye, Hyegang Jun, Deming ChenASPLOS 2024 · 被引用 21 次
- PipeLink: A Pipelined Resource Sharing System for Dataflow High-Level SynthesisRui Li, Lincoln Berkley, Rajit ManoharDAC 2025
