Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski
摘要
Datacenter network hotspots, defined as links with persistently high utilization, can lead to performance bottlenecks. In this work, we study hotspots in Google's datacenter networks. We find that these hotspots occur most frequently at ToR switches and can persist for hours. They are caused mainly by bandwidth demand-supply imbalance, largely due to high demand from network-intensive services, or demand exceeding available bandwidth when compute/storage upgrades outpace ToR bandwidth upgrades. Compounding this issue is bandwidth-independent task/data placement by datacenter compute and storage schedulers. We quantify the performance impact of hotspots, and find that they can degrade the end-to-end latency of some distributed applications by over 2× relative to low utilization levels. Finally, we describe simple improvements we deployed. In our cluster scheduler, adding hotspot-aware task placement reduced the number of hot ToRs by 90%; in our distributed file system, adding hotspot-aware data placement reduced p95 network latency by more than 50%. While congestion control, load balancing, and traffic engineering can efficiently utilize paths for a fixed placement, we find hotspot-aware placement -placing tasks and data under ToRs with higher available bandwidth -is crucial for achieving consistently good performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Borg: the next generationMuhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque 等EuroSys 2020 · 被引用 323 次
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh 等SIGCOMM 2022 · 被引用 230 次
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 被引用 144 次
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu 等SIGCOMM 2022 · 被引用 82 次
相关 Paper
- RDC: Energy-Efficient Data Center Network Congestion Relief with Topological Reconfigurability at the EdgeWeitao Wang, Dingming Wu, Sushovan Das, Afsaneh Rahbar 等NSDI 2022 · 被引用 14 次
- Fathom: Understanding Datacenter Application Network PerformanceMubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh 等SIGCOMM 2023 · 被引用 7 次
- Superways: A Datacenter Topology for Incast-heavy workloadsHamed Rezaei, Balajee VamananWWW 2021 · 被引用 6 次
- Saba: Rethinking Datacenter Network Allocation from Application's PerspectiveM. R. Siavash Katebzadeh, Paolo Costa, Boris GrotEuroSys 2023 · 被引用 4 次
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin 等EuroSys 2021 · 被引用 60 次
