TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
Neha Prakriya, Yuze Chi, Suhail Basalama, Linghao Song, Jason Cong
Abstract
Despite the increasing adoption of Field-Programmable Gate Arrays (FPGAs) in compute clouds, there remains a significant gap in programming tools and abstractions which can leverage network-connected, cloud-scale, multi-die FPGAs to generate accelerators with high frequency and throughput. To this end, we propose TAPA-CS, a task-parallel dataflow programming framework which automatically partitions and compiles a large design across a cluster of FPGAs with no additional user effort while achieving high frequency and throughput. TAPA-CS has three main contributions. First, it is an open-source framework which allows users to leverage virtually "unlimited" accelerator fabric, high-bandwidth memory (HBM), and on-chip memory, by abstracting away the underlying hardware. This reduces the user's programming burden to a logical one, enabling software developers and researchers with limited FPGA domain knowledge to deploy larger designs than possible earlier. Second, given as input a large design, TAPA-CS automatically partitions the design to map to multiple FPGAs, while ensuring congestion control, resource balancing, and overlapping of communication and computation. Third, TAPA-CS couples coarse-grained floorplanning with automated interconnect pipelining at the interand intra-FPGA levels to ensure high frequency. We have tested TAPA-CS on our multi-FPGA testbed where the FPGAs communicate through a high-speed 100Gbps Ethernet infrastructure. We have evaluated the performance and scalability of our tool on designs, including systolic-array based convolutional neural networks (CNNs), graph processing workloads such as page rank, stencil applications like the Dilate kernel, and K-nearest neighbors (KNN). TAPA-CS has the potential to accelerate development of increasingly complex and large designs on the low power and reconfigurable FPGAs. On average, the 2-FPGA, 3-FPGA, and 4-FPGA designs are 2.1×, 3.2×, and 4.4× faster than the single FPGA baselines generated through Vitis HLS. The superlinear speed-up is due to the increased memory bandwidth and parallelization opportunities enabled by multiple FPGAs. For each benchmark, we also analyze the computation to communication trade-offs that enable efficient multi-FPGA design. Through intelligent floorplanning and interconnect pipelining, TAPA-CS achieves a frequency improvement between 11%-116% compared with Vitis HLS. Lastly, we discuss the challenges in scaling to multiple server nodes containing 8 FPGAs, and directions for
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cc9319e-d66d-4a6d-87ba-9a9363f1da7fCited by top-tier papers1
Ask how each one uses itBuilds on5
- Virtualizing FPGAs in the CloudYue Zha, Jing LiASPLOS 2020 · 92 citations
- BYOC: A "Bring Your Own Core" Framework for Heterogeneous-ISA ResearchJonathan Balkind, Katie Lim, Michael Schaffner, Fei Gao et al.ASPLOS 2020 · 29 citations
- When application-specific ISA meets FPGAs: a multi-layer virtualization framework for heterogeneous cloud FPGAsYue Zha, Jing LiASPLOS 2021 · 24 citations
- Hetero-ViTAL: A Virtualization Stack for Heterogeneous FPGA ClustersYue Zha, Jing LiISCA 2021 · 21 citations
- SMAPPIC: Scalable Multi-FPGA Architecture Prototype Platform in the CloudGrigory Chirkov, David WentzlaffASPLOS 2023 · 13 citations
Related papers
- CODO: An Automated Compiler for Comprehensive Dataflow OptimizationWeichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang et al.ISCA 2026
- SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry ClusteringTianqi Zhang, Neha Prakriya, Sumukh Pinge, Jason Cong et al.DAC 2024 · 1 citation
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesZelin Du, Wei Zhang, Zimeng Zhou, Zili Shao et al.DAC 2023 · 9 citations
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
