Lune

SC2025顶会

A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC Workloads

Nicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni, Dhabaleswar K. Panda

2025年份
1被引次数

摘要

Dataflow accelerators can provide energy efficient and high-performance alternatives to current popular architectures. However, little work has been done to enable accelerator-initiated, scalable collective communication for these architectures. We develop a High Level Synthesis (HLS) interface to bridge this gap through software-hardware co-design. Given the tendency of dataflow applications to use reads and writes to streams to express data transfer, we develop a streaming interface implementing fine-grained transfers to the host processor. Data can then be communicated through MPI and transferred to the receiving accelerator. As a result, the interface uses few hardware resources for communication. As a case study, we enhance the HPL benchmark from HPCC_FPGA with our contributions. We evaluate our final design on up to 16 FPGAs, achieving up to 18% improvement in application throughput and 36% reduced latency in kernel execution. Additionally, we design a stencil benchmark that showcases superlinear speedup in a strong scaling scenario.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get a088050a-78b0-403f-8e9d-45e7a1bd03a0

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖