Saba: Rethinking Datacenter Network Allocation from Application's Perspective
M. R. Siavash Katebzadeh, Paolo Costa, Boris Grot
摘要
Today's datacenter workloads increasingly comprise distributed data-intensive applications, including data analytics, graph processing, and machine-learning training. These applications are bandwidth-hungry and often congest the datacenter network, resulting in poor network performance, which hurts application completion time. Efforts made to address this problem generally aim to achieve max-min fairness at the flow or application level. We observe that splitting the bandwidth equally among workloads is sub-optimal for aggregate application-level performance because various workloads exhibit different sensitivity to network bandwidth: for some workloads, even a small reduction in the available bandwidth yields a significant increase in completion time; for others, the completion time is largely insensitive to the available bandwidth.
Building on this insight, we propose Saba, an applicationaware bandwidth allocation framework that distributes network bandwidth based on application-level sensitivity. Saba combines ahead-of-time application profiling to determine bandwidth sensitivity with runtime bandwidth allocation using lightweight software support with no modifications to network hardware or protocols. Experiments with a 32server hardware testbed show that Saba improves average completion time by 1.88× (and by 1.27× in a simulated 1,944server cluster).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Taming unbalanced training workloads in deep learning with partial collective operationsShigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh 等PPoPP 2020 · 被引用 52 次
- Sol: Fast Distributed Computation Over Slow NetworksFan Lai, Jie You, Xiangfeng Zhu, Harsha V. Madhyastha 等NSDI 2020 · 被引用 24 次
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
相关 Paper
- Söze: One Network Telemetry Is All You Need for Per-flow Weighted Bandwidth Allocation at ScaleWeitao Wang, T. S. Eugene NgOSDI 2025 · 被引用 1 次
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu 等NSDI 2025 · 被引用 3 次
- Looking Beyond GPUs for DNN Scheduling on Multi-Tenant ClustersJayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, Vijay ChidambaramOSDI 2022 · 被引用 91 次
- Expanding across time to deliver bandwidth efficiency and low latencyWilliam M. Mellette, Rajdeep Das, Yibo Guo, Rob McGuinness 等NSDI 2020 · 被引用 194 次
- Scalable Real-Time Bandwidth Fairness in SwitchesRobert MacDavid, Xiaoqi Chen, Jennifer RexfordINFOCOM 2023 · 被引用 15 次
