CCEval: Accurately and Confidently Evaluating Performance Metrics of Congestion Control Algorithms for Datacenter Networks
Tianfeng Liu, Kaihui Gao, Li Chen, Dan Li, Jin Guang, Xinyun Chen, Vincent Liu, Zhiyong Chen, Yiwei Zhang, Ni Jin, Ran Zhang
Abstract
Congestion control in datacenter networks (DCNs) is a highly active research area. Typical CCA evaluation workflows contain three steps: generate experimental configurations, execute the experiments, and estimate performance metrics using results from multiple trials. However, due to variability brought by random traffic workloads and single-digit trial counts, common experimental methodologies fail to provide enough confidence to properly evaluate CCA performance.
We propose CCEval, an evaluation framework for accurately and confidently estimating performance metrics of CCAs in DCNs. The key idea is using confidence intervals and more trials to quantify and improve the accuracy and confidence of performance metrics. To this end, we propose a model-free estimation algorithm to calculate the confidence intervals and forecast the required trial count for a given accuracy, confidence level, metric, and CCA. We further design a model-based tail quantile estimation algorithm to reduce the needed trial counts significantly without losing accuracy and confidence. Extensive experiments on simulators and real-world testbeds with four CCAs on typical topologies and flow distributions show that CCEval can produce estimations of performance metrics accurately and confidently, with 1% relative margin of error and 95% confidence level, and can reduce trial counts by 75%∼80% for tail quantile estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00dc4b65-192d-4101-9eba-126845fa6299Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu et al.NSDI 2025 · 82 citations
- Is Big Data Performance Reproducible in Modern Cloud Networks?Alexandru Uta, Alexandru Custura, Dmitry Duplyakin, Ivo Jimenez et al.NSDI 2020 · 74 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
- DeepQueueNet: towards scalable and generalized network performance estimation with packet-level visibilityQingqing Yang, Xi Peng, Li Chen, Libin Liu et al.SIGCOMM 2022 · 43 citations
Related papers
- m3: Accurate Flow-Level Performance Estimation using Machine LearningChenning Li, Arash Nasr-Esfahany, Kevin Zhao, Kimia Noorbakhsh et al.SIGCOMM 2024 · 12 citations
- CClinguist: An Expert-Free Framework for Future-Compatible Congestion Control Algorithm IdentificationJiahui Li, Han Qi, Ruyi Yao, Jialin Wei et al.SIGCOMM 2025 · 2 citations
- Toward formally verifying congestion control behaviorVenkat Arun, Mina Tahmasbi Arashloo, Ahmed Saeed, Mohammad Alizadeh et al.SIGCOMM 2021 · 31 citations
- FRCC: Towards Provably Fair and Robust Congestion ControlAnup Agarwal, Venkat Arun, Srinivasan SeshanNSDI 2026
- CCAnalyzer: An Efficient and Nearly-Passive Congestion Control ClassifierRanysha Ware, Adithya Abraham Philip, Nicholas Hungria, Yash Kothari et al.SIGCOMM 2024 · 16 citations
