Orca: Server-assisted Multicast for Datacenter Networks
Khaled Diab, Parham Yassini, Mohamed Hefeeda
Abstract
Group communications appear in various large-scale datacenter applications. These applications, however, do not currently benefit from multicast, despite its potential substantial savings in network and processing resources. This is because current multicast systems do not scale and they impose considerable state and communication overheads. We propose a new architecture, called Orca, that addresses the challenges of multicast in datacenter networks. Orca divides the state and tasks of the data plane among switches and servers, and it partially offloads the management of multicast sessions to servers. Orca significantly reduces the state at switches, minimizes the bandwidth overhead, incurs small and constant processing overhead, and does not limit the size of multicast sessions. We implemented Orca in a testbed to demonstrate its performance in terms of throughput, consumption of server resources, packet latency, and the impact of server failures. We also implemented a sample multicast application in our testbed, and showed that Orca can substantially reduce its communication time, through optimizing the data transfer between nodes using multicast instead of unicast. In addition, we simulated a datacenter consisting of 27,648 hosts and handling 1M multicast sessions, and we compared Orca versus the state-of-art system in the literature. Our results show that Orca reduces the switch state by up to two orders of magnitude, the communication overhead by up to 19X, and the control overhead by up to 14X, compared to the state-of-art.
current datacenter multicast approaches, e.g., [5,6,32,43], improve upon the basic IP multicast, they also do not scale well and impose substantial overheads on the network, as we show in this paper. To partially mitigate the lack of efficient multicast systems, many datacenter applications had to rely on (inefficient) application-layer protocols [51]. For example, Apache Spark [40] implements its own primitives [9] such as Cornet [36] and HTTP-based multicast. This paper presents a new architecture, called Orca, to realize efficient multicast forwarding that can support millions of concurrent multicast sessions in datacenter networks. The idea of Orca is to offload some of the state maintained at network switches to end servers. To achieve this idea, Orca computes fixed-size and compact labels and attaches them to packets of multicast sessions. These labels effectively enable shifting some of the data plane tasks to servers. As a result, Orca significantly reduces the state at switches, minimizes the bandwidth overhead, incurs small and constant processing overhead, does not limit the size of multicast sessions, and eliminates redundant traffic. Realizing the proposed serverassisted multicast approach, however, faces multiple challenges at the control and data planes that Orca addresses. At the control plane, the proposed architecture needs to calculate optimized labels, manage state at servers, and handle failures. At the data plane, it requires packet processing algorithms at switches and servers that sustain the line-rate performance and minimize the latency and resource consumption.
This paper makes the following contributions.
• We introduce the idea of server-assisted (or offloaded) multicast for scalable multicast services in datacenters.
• We design a hierarchical control plane that efficiently manages multicast sessions and their dynamics, handles network failures, and does not impose high control overheads ( §3.3 and §3.4).
• We present a scalable data plane algorithm to process multicast packets within high-speed datacenter networks, without introducing redundant traffic or requiring switches to maintain large states ( §3.5).
• We design and implement APIs to transparently integrate multicast into datacenter applications ( §4).
• We implement the proposed multicast approach and evaluate its performance in a testbed using programmable switches to demonstrate its practicality ( §5). Our results show that the proposed approach can support high-speed traffic, uses small CPU resources at servers, and imposes small and predictable packet delays.
• We show the potential significant gains achieved by using multicast instead of unicast in datacenter applications. We implemented a sample application using Orca and the unicast approach used in current systems such as Apache Spark [40]. For this application that has only 12 receivers, our results show that Orca can reduce the communication time by almost an order of magnitude; larger savings are expected for applications with more receivers. In addition, since an Orca sender transmits only one copy per packet regardless of the number of receivers in the session, the required CPU resources are significantly reduced, compared to using unicast.
• We compare Orca against the closest system in the literature, Elmo [5], in large-scale simulations ( §6). Our results show that Orca reduces the switch state by up to two orders of m
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastWenxue Li, Junyi Zhang, Yufei Liu, Gaoxiong Zeng et al.HPCA 2024 · 13 citations
- Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and AggregationZhaoyi Li, Jiawei Huang, Yijun Li, Jingling Liu et al.USENIX ATC 2025 · 3 citations
Builds on2
Related papers
- SplitCast: Optimizing Multicast Flows in Reconfigurable Datacenter NetworksLong Luo, Klaus-Tycho Foerster, Stefan Schmid, Hongfang YuINFOCOM 2020 · 26 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
- Cloudcast: High-Throughput, Cost-Aware Overlay Multicast in the CloudSarah Wooders, Shu Liu, Paras Jain, Xiangxi Mo et al.NSDI 2024 · 20 citations
- Zeta: A Scalable and Robust East-West Communication Framework in Large-Scale CloudsQianyu Zhang, Gongming Zhao, Hongli Xu, Zhuolong Yu et al.NSDI 2022 · 10 citations
