Saba: Rethinking Datacenter Network Allocation from Application's Perspective
M. R. Siavash Katebzadeh, Paolo Costa, Boris Grot
Abstract
Today's datacenter workloads increasingly comprise distributed data-intensive applications, including data analytics, graph processing, and machine-learning training. These applications are bandwidth-hungry and often congest the datacenter network, resulting in poor network performance, which hurts application completion time. Efforts made to address this problem generally aim to achieve max-min fairness at the flow or application level. We observe that splitting the bandwidth equally among workloads is sub-optimal for aggregate application-level performance because various workloads exhibit different sensitivity to network bandwidth: for some workloads, even a small reduction in the available bandwidth yields a significant increase in completion time; for others, the completion time is largely insensitive to the available bandwidth.
Building on this insight, we propose Saba, an applicationaware bandwidth allocation framework that distributes network bandwidth based on application-level sensitivity. Saba combines ahead-of-time application profiling to determine bandwidth sensitivity with runtime bandwidth allocation using lightweight software support with no modifications to network hardware or protocols. Experiments with a 32server hardware testbed show that Saba improves average completion time by 1.88× (and by 1.27× in a simulated 1,944server cluster).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28c5c3cc-528d-4329-a43e-6df640138d85Builds on4
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh et al.PPoPP 2020 · 52 citations
- Sol: Fast Distributed Computation Over Slow NetworksFan Lai, Jie You, Xiangfeng Zhu, Harsha V. Madhyastha et al.NSDI 2020 · 24 citations
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
Related papers
- Söze: One Network Telemetry Is All You Need for Per-flow Weighted Bandwidth Allocation at ScaleWeitao Wang, T. S. Eugene NgOSDI 2025 · 1 citation
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu et al.NSDI 2025 · 3 citations
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 91 citations
- Expanding across time to deliver bandwidth efficiency and low latencyWilliam M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness et al.NSDI 2020 · 194 citations
- Scalable Real-Time Bandwidth Fairness in SwitchesRobert MacDavid, Xiaoqi Chen, Jennifer RexfordINFOCOM 2023 · 15 citations
