TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
Neha Prakriya, Yuze Chi, Suhail Basalama, Linghao Song, Jason Cong
摘要
Despite the increasing adoption of Field-Programmable Gate Arrays (FPGAs) in compute clouds, there remains a significant gap in programming tools and abstractions which can leverage network-connected, cloud-scale, multi-die FPGAs to generate accelerators with high frequency and throughput. To this end, we propose TAPA-CS, a task-parallel dataflow programming framework which automatically partitions and compiles a large design across a cluster of FPGAs with no additional user effort while achieving high frequency and throughput. TAPA-CS has three main contributions. First, it is an open-source framework which allows users to leverage virtually "unlimited" accelerator fabric, high-bandwidth memory (HBM), and on-chip memory, by abstracting away the underlying hardware. This reduces the user's programming burden to a logical one, enabling software developers and researchers with limited FPGA domain knowledge to deploy larger designs than possible earlier. Second, given as input a large design, TAPA-CS automatically partitions the design to map to multiple FPGAs, while ensuring congestion control, resource balancing, and overlapping of communication and computation. Third, TAPA-CS couples coarse-grained floorplanning with automated interconnect pipelining at the interand intra-FPGA levels to ensure high frequency. We have tested TAPA-CS on our multi-FPGA testbed where the FPGAs communicate through a high-speed 100Gbps Ethernet infrastructure. We have evaluated the performance and scalability of our tool on designs, including systolic-array based convolutional neural networks (CNNs), graph processing workloads such as page rank, stencil applications like the Dilate kernel, and K-nearest neighbors (KNN). TAPA-CS has the potential to accelerate development of increasingly complex and large designs on the low power and reconfigurable FPGAs. On average, the 2-FPGA, 3-FPGA, and 4-FPGA designs are 2.1×, 3.2×, and 4.4× faster than the single FPGA baselines generated through Vitis HLS. The superlinear speed-up is due to the increased memory bandwidth and parallelization opportunities enabled by multiple FPGAs. For each benchmark, we also analyze the computation to communication trade-offs that enable efficient multi-FPGA design. Through intelligent floorplanning and interconnect pipelining, TAPA-CS achieves a frequency improvement between 11%-116% compared with Vitis HLS. Lastly, we discuss the challenges in scaling to multiple server nodes containing 8 FPGAs, and directions for
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Virtualizing FPGAs in the CloudYue Zha, Jing LiASPLOS 2020 · 被引用 92 次
- BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA ResearchJonathan Balkind, Katie Lim, Michael Schaffner, Fei Gao 等ASPLOS 2020 · 被引用 29 次
- When application-specific ISA meets FPGAs: a multi-layer virtualization framework for heterogeneous cloud FPGAsYue Zha, Jing LiASPLOS 2021 · 被引用 24 次
- Hetero-ViTAL: A Virtualization Stack for Heterogeneous FPGA ClustersYue Zha, Jing LiISCA 2021 · 被引用 21 次
- SMAPPIC: Scalable Multi-FPGA Architecture Prototype Platform in the CloudGrigory Chirkov, David WentzlaffASPLOS 2023 · 被引用 13 次
相关 Paper
- CODO: An Automated Compiler for Comprehensive Dataflow OptimizationWeichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang 等ISCA 2026
- SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry ClusteringTianqi Zhang, Neha Prakriya, Sumukh Pinge, Jason Cong 等DAC 2024 · 被引用 1 次
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang 等HPCA 2026 · 被引用 2 次
- Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesZelin Du, Wei Zhang, Zimeng Zhou, Zili Shao 等DAC 2023 · 被引用 9 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
