Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable Multicast
Wenxue Li, Junyi Zhang, Yufei Liu, Gaoxiong Zeng, Zilong Wang, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen
摘要
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang 等SIGCOMM 2025 · 被引用 10 次
- Enabling In-Network Acceleration Over the CloudHao Wang, Decang Sun, Jinbin Hu, Kai ChenINFOCOM 2025 · 被引用 3 次
- A Generic and Efficient Communication Framework for Message-Level In-Network ComputingXinchen Wan, Luyang Li, Han Tian, Xudong Liao 等INFOCOM 2025 · 被引用 2 次
- EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou 等SIGCOMM 2026 · 被引用 1 次
- RDNet: An RDMA-aware Container Network Interface for Cloud EnvironmentsMyoungsung You, Minjae Seo, Seungwon Shin, Jaehyun NamINFOCOM 2026
它引用的顶会 Paper13
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi 等NSDI 2021 · 被引用 228 次
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng 等NSDI 2023 · 被引用 154 次
- From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudRui Miao, Lingjun Zhu, Shu Ma, Kun Qian 等SIGCOMM 2022 · 被引用 79 次
相关 Paper
- Orca: Server-assisted Multicast for Datacenter NetworksKhaled Diab, Parham Yassini, Mohamed HefeedaNSDI 2022 · 被引用 13 次
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- EDM: An Ultra-Low Latency Ethernet Fabric for Memory DisaggregationWeigao Su, Vishal ShrivastavASPLOS 2025 · 被引用 1 次
- Scalable RDMA Transport with Efficient Connection SharingJian Tang, Xiaoliang Wang, Huichen DaiINFOCOM 2023 · 被引用 13 次
- Network Load Balancing with In-network Reordering Support for RDMACha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi 等SIGCOMM 2023 · 被引用 110 次
