RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICs
Maximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu, Zhiyi Hu, Jonas Dann, Gustavo Alonso
摘要
Remote Direct Memory Access (RDMA) has become the de facto standard for high-performance data center networking. However, current deployments rely heavily on fixed-function, commercial NICs. These "black box" commercial hardware implementations prevent researchers and system architects from modifying the transport layer for specialized tasks. In parallel, research on NICs often lacks offloaded networking stacks or uses simplified protocol implementations, limiting insight into novel networking solutions in realistic settings. In this paper, we bridge this gap by introducing BALBOA, an open-source, 100 Gbps RDMA offload engine designed for research on networking and fully compatible with commercial RNICs. Unlike prior stack implementations which lack scalability and bandwidth, or struggle with data center interoperability and miss strict protocol compliance, BAL-BOA supports hundreds of Queue Pairs in switched network environments and allows for line-rate offloads, making it a viable platform for realistic data center research. We describe the system architecture, detailing how BALBOA overcomes FPGA memory and timing bottlenecks through a decoupled state architecture and streaming control-data separation. We evaluate BALBOA on a hardware cluster with FPGAs, RNICs, and switches, showing that it matches the performance of commercial ASICs while offering full customization. Finally, we showcase BALBOA's potential through novel case studies: protocol enhancements for infrastructure purposes (encryption, deep packet inspection) and an offloaded preprocessing pipeline for deep learning recommender systems, which applies streaming transformations to the incoming data before feeding it directly to a GPU for model serving.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu 等SIGCOMM 2024 · 被引用 171 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- Do OS abstractions make sense on FPGAs?Dario Korolija, Timothy Roscoe, Gustavo AlonsoOSDI 2020 · 被引用 114 次
- One-sided RDMA-Conscious Extendible Hashing for Disaggregated MemoryPengfei Zuo, Jiazhao Sun, Liu Yang, Shuangwu Zhang 等USENIX ATC 2021 · 被引用 113 次
相关 Paper
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersArjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch 等SIGCOMM 2020 · 被引用 70 次
- Network Load Balancing with In-network Reordering Support for RDMACha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi 等SIGCOMM 2023 · 被引用 110 次
- Tlaloc: A Generic Multipath Load Balancing for RoCEHuimin Luo, Jiao Zhang, Yongchen Pan, Tian Pan 等INFOCOM 2026
- StRoM: smart remote memoryDavid Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni 等EuroSys 2020 · 被引用 83 次
- ACCL+: an FPGA-Based Collective Engine for Distributed ApplicationsZhenhao He, Dario Korolija, Yu Zhu, Benjamin Ramhorst 等OSDI 2024 · 被引用 12 次
