Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and Storage
Hamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu, Ramesh Govindan, Neal Cardwell, Nandita Dukkipati, Meng-Jung Tsai, Chris DeForeest, Yuxue Jin, Charles J. Carver, Jan Kopanski
Abstract
Datacenter network hotspots, defined as links with persistently high utilization, can lead to performance bottlenecks. In this work, we study hotspots in Google's datacenter networks. We find that these hotspots occur most frequently at ToR switches and can persist for hours. They are caused mainly by bandwidth demand-supply imbalance, largely due to high demand from network-intensive services, or demand exceeding available bandwidth when compute/storage upgrades outpace ToR bandwidth upgrades. Compounding this issue is bandwidth-independent task/data placement by datacenter compute and storage schedulers. We quantify the performance impact of hotspots, and find that they can degrade the end-to-end latency of some distributed applications by over 2× relative to low utilization levels. Finally, we describe simple improvements we deployed. In our cluster scheduler, adding hotspot-aware task placement reduced the number of hot ToRs by 90%; in our distributed file system, adding hotspot-aware data placement reduced p95 network latency by more than 50%. While congestion control, load balancing, and traffic engineering can efficiently utilize paths for a fixed placement, we find hotspot-aware placement -placing tasks and data under ToRs with higher available bandwidth -is crucial for achieving consistently good performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext beb819bc-1823-4efa-be8f-d082bc82d422Builds on13
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- Borg: the next generationMuhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque et al.EuroSys 2020 · 323 citations
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh et al.SIGCOMM 2022 · 230 citations
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 144 citations
- PLB: congestion signals are simple and effective for network load balancingMubashir Adnan Qureshi, Yuchung Cheng, Qianwen Yin, Qiaobin Fu et al.SIGCOMM 2022 · 82 citations
Related papers
- RDC: Energy-Efficient Data Center Network Congestion Relief with Topological Reconfigurability at the EdgeWeitao Wang, Dingming Wu, Sushovan Das, Afsaneh Rahbar et al.NSDI 2022 · 14 citations
- Fathom: Understanding Datacenter Application Network PerformanceMubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh et al.SIGCOMM 2023 · 7 citations
- Superways: A Datacenter Topology for Incast-heavy workloadsHamed Rezaei, Balajee VamananWWW 2021 · 6 citations
- Saba: Rethinking Datacenter Network Allocation from Application's PerspectiveM. R. Siavash Katebzadeh, Paolo Costa, Boris GrotEuroSys 2023 · 4 citations
- Take it to the limit: peak prediction-driven resource overcommitment in datacentersNoman Bashir, Nan Deng, Krzysztof Rzadca, David Irwin et al.EuroSys 2021 · 60 citations
