Scalable Tail Latency Estimation for Data Center Networks
Kevin Zhao, Prateesh Goyal, Mohammad Alizadeh, Thomas E. Anderson
摘要
In this paper, we consider how to provide fast estimates of flow-level tail latency performance for very large scale data center networks. Network tail latency is often a crucial metric for cloud application performance that can be affected by a wide variety of factors, including network load, inter-rack traffic skew, traffic burstiness, flow size distributions, oversubscription, and topology asymmetry. Network simulators such as ns-3 and OMNeT++ can provide accurate answers, but are very hard to parallelize, taking hours or days to answer what if questions for a single configuration at even moderate scale. Recent work with MimicNet has shown how to use machine learning to improve simulation performance, but at a cost of including a long training step per configuration, and with assumptions about workload and topology uniformity that typically do not hold in practice. We address this gap by developing a set of techniques to provide fast performance estimates for large scale networks with general traffic matrices and topologies. A key step is to decompose the problem into a large number of parallel independent single-link simulations; we carefully combine these link-level simulations to produce accurate estimates of end-to-end flow level performance distributions for the entire network. Like MimicNet, we exploit symmetry where possible to gain additional speedups, but without relying on machine learning, so there is no training delay. On large-scale networks where ns-3 takes 11 to 27 hours to simulate five seconds of network behavior, our techniques run in one to two minutes with 99th percentile accuracy within 9% for flow completion times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Unison: A Parallel-Efficient and User-Transparent Network Simulation KernelSongyuan Bai, Hao Zheng, Chen Tian, Xiaoliang Wang 等EuroSys 2024 · 被引用 20 次
- SketchPolymer: Estimate Per-item Tail Quantile Using One SketchJiarui Guo, Yisen Hong, Yuhan Wu, Yunfei Liu 等KDD 2023 · 被引用 13 次
- m3: Accurate Flow-Level Performance Estimation using Machine LearningChenning Li, Arash Nasr-Esfahany, Kevin Zhao, Kimia Noorbakhsh 等SIGCOMM 2024 · 被引用 12 次
- Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-ForwardingFei Long, Kaihui Gao, Li Chen, Dan Li 等NSDI 2026 · 被引用 4 次
- CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter NetworksTianfeng Liu, Kaihui Gao, Li Chen, Dan Li 等NSDI 2026 · 被引用 1 次
它引用的顶会 Paper4
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- SP-PIFO: Approximating Push-In First-Out Behaviors using Strict-Priority QueuesAlbert Gran Alcoz, Alexander Dietmüller, Laurent VanbeverNSDI 2020 · 被引用 140 次
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan 等SIGCOMM 2021 · 被引用 63 次
- DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibilityQingqing Yang, Xi Peng, Li Chen, Libin Liu 等SIGCOMM 2022 · 被引用 43 次
相关 Paper
- DONS: Fast and Affordable Discrete Event Network Simulation with Automatic ParallelizationKaihui Gao, Li Chen, Dan Li, Vincent Liu 等SIGCOMM 2023 · 被引用 30 次
- Optimizing Network Simulation: Enhancing Performance Prediction Accuracy via Neural Architecture SearchShaoChen He, Zirui Zhuang, Haifeng Sun, Xiaoyuan Fu 等ICML 2026
- On Modular Learning of Distributed Systems for Predicting End-to-End LatencyChieh-Jan Mike Liang, Zilin Fang, Yuqing Xie, Fan Yang 等NSDI 2023 · 被引用 19 次
- Days: Discrete-Event Network Simulation on SteroidsBaochun LiINFOCOM 2026
- Towards Domain-Specific Network Transport for Distributed DNN TrainingHao Wang, Han Tian, Jingrong Chen, Xinchen Wan 等NSDI 2024 · 被引用 54 次
