Credence: Augmenting Datacenter Switch Buffer Sharing with ML Predictions
Vamsi Addanki, Maciej Pacut, Stefan Schmid
摘要
Packet buffers in datacenter switches are shared across all the switch ports in order to improve the overall throughput. The trend of shrinking buffer sizes in datacenter switches makes buffer sharing extremely challenging and a critical performance issue. Literature suggests that push-out buffer sharing algorithms have significantly better performance guarantees compared to drop-tail algorithms. Unfortunately, switches are unable to benefit from these algorithms due to lack of support for push-out operations in hardware. Our key observation is that drop-tail buffers can emulate push-out buffers if the future packet arrivals are known ahead of time. This suggests that augmenting drop-tail algorithms with predictions about the future arrivals has the potential to significantly improve performance. This paper is the first research attempt in this direction. We propose Credence, a drop-tail buffer sharing algorithm augmented with machine-learned predictions. Credence can unlock the performance only attainable by push-out algorithms so far. Its performance hinges on the accuracy of predictions. Specifically, Credence achieves near-optimal performance of the best known push-out algorithm LQD (Longest Queue Drop) with perfect predictions, but gracefully degrades to the performance of the simplest drop-tail algorithm Complete Sharing when the prediction error gets arbitrarily worse. Our evaluations show that Credence improves throughput by x compared to traditional approaches. In terms of flow completion times, we show that Credence improves upon the state-of-the-art approaches by up to using off-the-shelf machine learning techniques that are also practical in today's hardware. We believe this work opens several interesting future work opportunities both in systems and theory that we discuss at the end of this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Leo: Online ML-based Traffic Classification at Multi-Terabit Line RateSyed Usman Jafri, Sanjay G. Rao, Vishal Shrivastav, Mohit TawarmalaniNSDI 2024 · 被引用 46 次
- Reverie: Low Pass Filter-Based Switch Buffer Sharing for Datacenters with RDMA and TCP TrafficVamsi Addanki, Wei Bai, Stefan Schmid, Maria ApostolakiNSDI 2024 · 被引用 33 次
- Occamy: A Preemptive Buffer Management for On-chip Shared-memory SwitchesDanfeng Shan, Yunguang Li, Jinchao Ma, Zhenxing Zhang 等EuroSys 2025 · 被引用 5 次
- Themis: Scheduling-Aware Buffer Management for HBM-Based Hybrid BuffersZhiyu Zhang, Minkun Xue, Kan Yu, Ruyi Yao 等NSDI 2026
它引用的顶会 Paper14
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- PowerTCP: Pushing the Performance Limits of Datacenter NetworksVamsi Addanki, Oliver Michel, Stefan SchmidNSDI 2022 · 被引用 116 次
- Annulus: A Dual Congestion Control Loop for Datacenter and WAN Traffic AggregatesAhmed Saeed, Varun Gupta, Prateesh Goyal, Milad Sharif 等SIGCOMM 2020 · 被引用 72 次
相关 Paper
- OBM: Optimal Shared Packet Buffer Management in SwitchesDan Mani Binu, Jason Lei, Vishal ShrivastavSIGCOMM 2026
- One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersWei Bai, Shuihai Hu, Kai Chen, Kun Tan 等INFOCOM 2020 · 被引用 34 次
- Learning Buffer Management Policies for Shared Memory SwitchesMowei Wang, Sijiang Huang, Yong Cui, Wendong Wang 等INFOCOM 2022 · 被引用 12 次
- Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in DatacentersSong Zhang, Wenxin Li, Yulong Li, Yuan Liu 等INFOCOM 2025
- Traffic-aware Buffer Management in Shared Memory SwitchesSijiang Huang, Mowei Wang, Yong CuiINFOCOM 2021 · 被引用 14 次
