CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen, Dan Li, Jin Guang, Xinyun Chen, Vincent Liu, Zhiyong Chen, Yiwei Zhang, Ni Jin, Ran Zhang
摘要
Congestion control in datacenter networks (DCNs) is a highly active research area. Typical CCA evaluation workflows contain three steps: generate experimental configurations, execute the experiments, and estimate performance metrics using results from multiple trials. However, due to variability brought by random traffic workloads and single-digit trial counts, common experimental methodologies fail to provide enough confidence to properly evaluate CCA performance.
We propose CCEval, an evaluation framework for accurately and confidently estimating performance metrics of CCAs in DCNs. The key idea is using confidence intervals and more trials to quantify and improve the accuracy and confidence of performance metrics. To this end, we propose a model-free estimation algorithm to calculate the confidence intervals and forecast the required trial count for a given accuracy, confidence level, metric, and CCA. We further design a model-based tail quantile estimation algorithm to reduce the needed trial counts significantly without losing accuracy and confidence. Extensive experiments on simulators and real-world testbeds with four CCAs on typical topologies and flow distributions show that CCEval can produce estimations of performance metrics accurately and confidently, with 1% relative margin of error and 95% confidence level, and can reduce trial counts by 75%∼80% for tail quantile estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- Is Big Data Performance Reproducible in Modern Cloud Networks?Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez 等NSDI 2020 · 被引用 74 次
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan 等SIGCOMM 2021 · 被引用 63 次
- DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibilityQingqing Yang, Xi Peng, Li Chen, Libin Liu 等SIGCOMM 2022 · 被引用 43 次
相关 Paper
- m3: Accurate Flow-Level Performance Estimation using Machine LearningChenning Li, Arash Nasr-Esfahany, Kevin Zhao, Kimia Noorbakhsh 等SIGCOMM 2024 · 被引用 12 次
- CClinguist: An Expert-Free Framework for Future-Compatible Congestion Control Algorithm IdentificationJiahui Li, Han Qi, Ruyi Yao, Jialin Wei 等SIGCOMM 2025 · 被引用 2 次
- Toward formally verifying congestion control behaviorVenkat Arun, Mina Tahmasbi Arashloo, Ahmed Saeed, Mohammad Alizadeh 等SIGCOMM 2021 · 被引用 31 次
- FRCC: Towards Provably Fair and Robust Congestion ControlAnup Agarwal, Venkat Arun, Srinivasan SeshanNSDI 2026
- CCAnalyzer: An Efficient and Nearly-Passive Congestion Control ClassifierRanysha Ware, Adithya Abraham Philip, Nicholas Hungria, Yash Kothari 等SIGCOMM 2024 · 被引用 16 次
