OSCAR: O(1)-Step Convergence and Readily-deployable Congestion Control
Zhaochen Zhang, Feiyang Xue, Rui Ning, Keqiang He, Gianni Antichi, Jiaqi Gao, Zhimeng Yin, Kexin Liu, Rui Li, Zhengqi Cui, Zhehao Lin, Peirui Cao
Abstract
Datacenter CCs typically target full bandwidth utilization and minimal queueing delay and strive to converge to these targets as quickly as possible. State-of-the-art CCs exhibit different convergence speeds, with the fastest ones converging in O(1) steps, which means reaching the target in constant time regardless of network conditions. However, their reliance on network features makes them not readily deployable. For instance, precise-INT-based CCs, such as HPCC and PowerTCP, achieve O(1)-step convergence through MIMD operations based on precise congestion information from the lengthy INT header, which is challenging to support for high-speed commodity hardware. Our key insight is that delay and delay gradient can exhibit precision comparable to INT, enabling O(1)-step convergence without specialized network features. Based on this insight, we propose OSCAR, the first O(1)-Step Convergence And Readily-deployable CC. OSCAR introduces novel techniques to accurately estimate the delay gradient with minimal overhead, eliminate overreaction in MIMD updates, and coordinate independent control loops to converge to one target. Testbed evaluations demonstrate OSCAR can rapidly converge to the fair share under real-world noise. In large-scale simulations with realistic workloads, OSCAR consistently outperforms precise-INT-based CCs by 12%-48% on average FCT and 40%-74% on tail FCT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d3d8f00-2f60-4743-b17a-308f6070f5ffBuilds on25
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui et al.FAST 2025 · 337 citations
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi et al.SIGCOMM 2020 · 268 citations
- When Cloud Storage Meets RDMAYixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi et al.NSDI 2021 · 228 citations
Related papers
- PowerTCP: Pushing the Performance Limits of Datacenter NetworksVamsi Addanki, Oliver Michel, Stefan SchmidNSDI 2022 · 116 citations
- Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INTWeitao Wang, Masoud Moshref, Yuliang Li, Gautam Kumar et al.NSDI 2023 · 58 citations
- CCC: Re-architecting Delay-based Congestion Control in Datacenter NetworksWanchun Jiang, Haoyang Li, Kai Wang, Yujie Hu et al.NSDI 2026 · 1 citation
- Simplifying Prioritization and Scheduling with P2CSAli Munir, Xiaolin Pang, Junyi ZhangSIGCOMM 2026
- ACC: Addressing Performance Limitations in Datacenters with Atomic Congestion ControlZirui Wan, Jiao Zhang, Shuai Cheng, Ying Chen et al.INFOCOM 2025 · 3 citations
