RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAs
Marco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer, Lucian Petrica, Dario Korolija, Marco D. Santambrogio, Davide Conficconi, Gustavo Alonso, Kenneth O'Brien
摘要
Modern datacenters integrate heterogeneous accelerators, such as GPUs and FPGAs, to speed up different stages of compute-intensive pipelines. GPUs are best suited for massively parallel workloads (e.g., deep learning), while FPGAs excel at task-level parallelism, stream-oriented processing, and in-network acceleration. Since these architectures must exchange data efficiently, literature introduced Peer-To-Peer (P2P) communication across PCI Express (PCIe) devices, to reduce CPU-driven orchestration and avoid intermediate, redundant buffer copies that degrade performance. However, current solutions are either closed-source or tied to proprietary frameworks, limiting P2P communication across most PCIe-based devices and requiring significant technical effort to enable P2P capabilities on supported hardware. For this reason, we propose RoPeerTo, a fully open-source, datacenter-scale architecture for P2P DMA communication over PCIe, validated on both GPUs and FPGAs. The goal is to provide a general, open alternative that ensures flexibility, efficiency, and usability. To this end, we design a complete HW/SW stack operating across different layers, supporting standard protocols for DMA-based memory sharing, advanced tools for device virtualization, memory address translation, and access protection. The result is a unified framework exposing a high-level API to end users, that enables direct communication between accelerators such as FPGAs and GPUs, and abstracts away the underlying hardware setup and management. We validate the system across different scenarios. First, we isolate the communication layer, observing a 5.61× speedup and a 37.99% reduction in GPU power consumption during data transfer. Next, we leverage the system for a compute-intensive workload where communication is only a partial bottleneck, achieving a 6.77% speedup without any compute-side modifications. Finally, we evaluate communication-heavy distributed computing workloads, demonstrating up to a 21.79× speedup in network-bound data scattering.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li 等ISCA 2026 · 被引用 1 次
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu 等OSDI 2026
相关 Paper
- Efficient Multi-GPU Shared Memory via Automatic Optimization of Fine-Grained TransfersHarini Muthukrishnan, David W. Nellans, Daniel Lustig, Jeffrey A. Fessler 等ISCA 2021 · 被引用 16 次
- ARK: GPU-driven Code Execution for Distributed Deep LearningChangho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu 等NSDI 2023 · 被引用 22 次
- ACCL+: an FPGA-Based Collective Engine for Distributed ApplicationsZhenhao He, Dario Korolija, Yu Zhu, Benjamin Ramhorst 等OSDI 2024 · 被引用 12 次
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU ClustersAntian Liang, Zhigang Zhao, Kai Zhang, Xuri Shi 等EuroSys 2026 · 被引用 1 次
- GPU-Ether: GPU-native Packet I/O for GPU Applications on Commodity EthernetChangue Jung, Suhwan Kim, Ikjun Yeom, Honguk Woo 等INFOCOM 2021 · 被引用 7 次
