ACCL+: an FPGA-Based Collective Engine for Distributed Applications
Zhenhao He, Dario Korolija, Yu Zhu, Benjamin Ramhorst, Tristan Laan, Lucian Petrica, Michaela Blott, Gustavo Alonso
摘要
FPGAs are increasingly prevalent in cloud deployments, serving as Smart NICs or network-attached accelerators. Despite their potential, developing distributed FPGA-accelerated applications remains cumbersome due to the lack of appropriate infrastructure and communication abstractions. To facilitate the development of distributed applications with FP-GAs, in this paper we propose ACCL+, an open-source versatile FPGA-based collective communication library. Portable across different platforms and supporting UDP, TCP, as well as RDMA, ACCL+ empowers FPGA applications to initiate direct FPGA-to-FPGA collective communication. Additionally, it can serve as a collective offload engine for CPU applications, freeing the CPU from networking tasks. It is user-extensible, allowing new collectives to be implemented and deployed without having to re-synthesize the FPGA circuit. We evaluated ACCL+ on an FPGA cluster with 100 Gb/s networking, comparing its performance against software MPI over RDMA. The results demonstrate ACCL+'s significant advantages for FPGA-based distributed applications and highly competitive performance for CPU applications. We showcase ACCL+'s dual role with two use cases: seamlessly integrating as a collective offload engine to distribute CPU-based vector-matrix multiplication, and serving as a crucial and efficient component in designing fully FPGA-based distributed deep-learning recommendation inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
- Coyote v2: Raising the Level of Abstraction for Data Center FPGAsBenjamin Ramhorst, Dario Korolija, Maximilian Jakob Heer, Jonas Dann 等SOSP 2025 · 被引用 6 次
- SwiftSpatial: Spatial Joins on Modern HardwareWenqi Jiang, Oleh-Yevhen Khavrona, Martin Parvanov, Gustavo AlonsoSIGMOD 2025 · 被引用 2 次
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min 等SIGCOMM 2026 · 被引用 1 次
它引用的顶会 Paper17
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
- Rethinking software runtimes for disaggregated memoryIrina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap 等ASPLOS 2021 · 被引用 116 次
- Do OS abstractions make sense on FPGAs?Dario Korolija, Timothy Roscoe, Gustavo AlonsoOSDI 2020 · 被引用 114 次
- Clio: a hardware-software co-designed disaggregated memory systemZhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang 等ASPLOS 2022 · 被引用 110 次
- PANIC: A High-Performance Programmable NIC for Multi-tenant NetworksJiaxin Lin, Kiran Patel, Brent E. Stephens, Anirudh Sivaraman 等OSDI 2020 · 被引用 104 次
相关 Paper
- TCCL: Discovering Better Communication Paths for PCIe GPU ClustersHeehoon Kim, Junyeol Ryu, Jaejin LeeASPLOS 2024 · 被引用 26 次
- UCCL-Tran: An Extensible Software Transport Layer for GPU NetworkingYang Zhou, Zhongjie Chen, Ziming Mao, ChonLam Lao 等OSDI 2026
- MSCCL++: Rethinking GPU Communication Abstractions for AI InferenceChangho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda 等ASPLOS 2026 · 被引用 5 次
- Enabling Efficient GPU Communication over Multiple NICs with FuseLinkZhenghang Ren, Yuxuan Li, Zilong Wang, Xinyang Huang 等OSDI 2025 · 被引用 10 次
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu 等OSDI 2026
