Flow Scheduling with Imprecise Knowledge
Wenxin Li, Xin He, Yuan Liu, Keqiu Li, Kai Chen, Zhao Ge, Zewei Guan, Heng Qi, Song Zhang, Guyue Liu
Abstract
Most existing data center network (DCN) flow scheduling solutions aim to minimize flow completion times (FCT). However, these solutions either require precise flow information (e.g., per-flow size), which is challenging to implement on commodity switches (e.g., pFabric [7]), or no prior flow information at all, which is at the cost of performance (e.g., PIAS [10]). In this work, we present QCLIMB, a new flow scheduling solution designed to minimize FCT by utilizing imprecise flow information. Our key observation is that although obtaining precise flow information can be challenging, it is possible to accurately estimate each flow's lower and upper bounds with machine learning techniques.
QCLIMB has two key parts: i) a novel scheduling algorithm that leverages the lower bounds of different flows to prioritize small flow over large flows from the beginning of transmission, rather than at later stages; and ii) an efficient out-of-order handling mechanism that addresses practical reordering issues resulting from the algorithm. We show that QCLIMB significantly outperforms PIAS (88% lower average FCT of small flows) and is surprisingly close to pFabric (around 9% gap) while not requiring any switch modifications. Schemes Requiring no switch changes or advanced hardware Using limited number of priority queues Retaining existing TCP/IP network stacks Using lower bounds for flow scheduling Clairvoyant
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang et al.NSDI 2025 · 12 citations
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang et al.SIGCOMM 2025 · 10 citations
- PPT: A Pragmatic Transport for DatacentersLide Suo, Yiren Pang, Wenxin Li, Renjie Pei et al.SIGCOMM 2024 · 7 citations
- Enabling Virtual Priority in Data Center Congestion ControlZhaochen Zhang, Feiyang Xue, Keqiang He, Zhimeng Yin et al.EuroSys 2025 · 6 citations
Builds on4
- Aeolus: A Building Block for Proactive Transport in DatacentersShuihai Hu, Wei Bai, Gaoxiong Zeng, Zilong Wang et al.SIGCOMM 2020 · 138 citations
- A Computational Approach to Packet ClassificationAlon Rashelbach, Ori Rottenstreich, Mark SilbersteinSIGCOMM 2020 · 65 citations
- One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersWei Bai, Shuihai Hu, Kai Chen, Kun Tan et al.INFOCOM 2020 · 34 citations
- dcPIM: near-optimal proactive datacenter transportQizhe Cai, Mina Tahmasbi Arashloo, Rachit AgarwalSIGCOMM 2022 · 30 citations
Related papers
- QCluster: Clustering Packets for Flow SchedulingTong Yang, Jizhou Li, Yikai Zhao, Kaicheng Yang et al.WWW 2022 · 11 citations
- Cutting Tail Latency in Commodity Datacenters with CloudburstGaoxiong Zeng, Li Chen, Bairen Yi, Kai ChenINFOCOM 2022 · 12 citations
- FLB: Fine-grained Load Balancing for Lossless Datacenter NetworksJinbin Hu, Wenxue Li, Xiangzhou Liu, Junfeng Wang et al.USENIX ATC 2025 · 10 citations
- Towards timeout-less transport in commodity datacenter networksHwijoon Lim, Wei Bai, Yibo Zhu, Youngmok Jung et al.EuroSys 2021 · 28 citations
- Credence: Augmenting Datacenter Switch Buffer Sharing with ML PredictionsVamsi Addanki, Maciej Pacut, Stefan SchmidNSDI 2024 · 22 citations
