Starfish: A Topology-Routing Co-Design for Small-Scale Data Centers
Anchengcheng Zhou, Vipul Harsh, Sangeetha Abdu Jyothi, Maria Apostolaki
Abstract
Most data centers today operate at scales much smaller than hyperscalers, i.e., hosting only hundreds to a few thousand servers, yet their design has received disproportionately little attention in the literature. Leaf-spine is nearly ubiquitous in small-scale data centers, but presents a fundamental tradeoff: provisioning full non-blocking capacity is costly, while oversubscription reduces cost but under-serves bursty, skewed workloads. Recent expander-based designs could, in theory, offer better performance but face substantial deployment hurdles, including lacking a practical routing scheme.
We introduce Starfish, a topology-routing co-design tailored to small-scale data centers, that delivers high performance, fault tolerance, and ease of deployment. Starfish's topology, DRing, increases per-server egress capacity by spreading servers across all switches, reduces latency-even under failure scenarios-by exposing diverse near-shortest paths, and eases deployment by organizing switches into uniform blocks. Starfish's routing leverages structural properties of DRing and effectively adapts to traffic patterns using standard hardware and protocols. Notably, Starfish's routing also generalizes to expander-based topologies.
Evaluation shows that Starfish supports 56% more traffic on average than leaf-spine built with the same equipment across real-world traces. This demonstrates that small-scale data centers today can realistically achieve higher performance with their existing equipment. Starfish's routing also enables an expander-based design to support 52% more traffic on average than leaf-spine. More broadly, this work takes a first step towards a distinct design space for small-scale data centers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh et al.SIGCOMM 2022 · 230 citations
- Expanding across time to deliver bandwidth efficiency and low latencyWilliam M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness et al.NSDI 2020 · 194 citations
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth et al.SC 2020 · 122 citations
- DOTE: Rethinking (Predictive) WAN Traffic EngineeringYarin Perry, Felipe Vieira Frujeri, Chaim Hoch, Srikanth Kandula et al.NSDI 2023 · 81 citations
- Cost-effective Cloud Edge Traffic Engineering with CascaraRachee Singh, Sharad Agarwal, Matt Calder, Paramvir BahlNSDI 2021 · 79 citations
Related papers
- A throughput-centric view of the performance of datacenter topologiesPooria Namyar, Sucha Supittayapornpong, Mingyang Zhang, Minlan Yu et al.SIGCOMM 2021 · 29 citations
- A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly NetworkNils Blach, Maciej Besta, Daniele De Sensi, Jens Domke et al.NSDI 2024 · 13 citations
- Orca: Server-assisted Multicast for Datacenter NetworksKhaled Diab, Parham Yassini, Mohamed HefeedaNSDI 2022 · 13 citations
- RDC: Energy-Efficient Data Center Network Congestion Relief with Topological Reconfigurability at the EdgeWeitao Wang, Dingming Wu, Sushovan Das, Afsaneh Rahbar et al.NSDI 2022 · 14 citations
- Beyond the mega-data center: networking multi-data center regionsVojislav Dukic, Ginni Khanna, Christos Gkantsidis, Thomas Karagiannis et al.SIGCOMM 2020 · 49 citations
