Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen, Xin Zhang, Can Liu, Hongyu Wang, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang, Yongfeng Liu, Zhigang Chen
Abstract
RDMA, with its high throughput, ultra-low latency, and low CPU utilization, has been widely used in large-scale data centers. However, due to the limited RDMA RNIC on-chip memory, commodity RDMA usually implements the simple Go-Back-N (GBN) loss recovery mechanism, which consumes less memory but leads to a significant performance drop when encountering loss. Recent works try to improve RDMA performance under packet loss by introducing selective retransmission (SR) to it. Nevertheless, implementing efficient SR in RDMA remains challenging. Specifically, either it consumes too much memory for maintaining SR states which leads to poor connection scalability, or it incurs high CPU consumption and latency for onloading SR processing back to the CPU software. To this end, we propose FaSR, a fast and scalable RDMA selective retransmission. It is fast by processing the SR with 200Gbps+ line-rate fully on the RDMA NIC chip, and is scalable by introducing novel SR state management schemes thus consuming small memory even under high concurrency. Specifically, FaSR adopts a dynamically sharing SR structure among connections to reduce the memory footprint by orders of magnitude when the concurrency is high. Also, utilizing the loss recovery pattern, FaSR devises several techniques thus it can access the sharing structure with line-rate for different connections. We have implemented FaSR in Xilinx FPGA board with 4000 lines of verilog code. Testbed evaluation demonstrates that FaSR can maintain 92%+ throughput at a packet loss rate of 1% under more than 5K concurrent connections, which is 16% and 12.6x higher compared to the latest RDMA SR solution and commodity RDMA NICs, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get dee762db-902b-4e3f-ae25-b8b597429e0aCited by top-tier papers3
- LCMP: Distributed Long-Haul Cost-Aware Multi-Path Routing for Inter-Datacenter RDMA NetworksDong-Yang Yu, Yuchao Zhang, Xiaodi Wang, Jun Wang et al.EuroSys 2026 · 2 citations
- LR2: Accelerating Long-Distance RDMA Recovery via In-Network Retransmission DecouplingMinfei Long, Jiangping Han, Kaiping Xue, Jinhao Liu et al.INFOCOM 2026
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu et al.OSDI 2026
Related papers
- SRNIC: A Scalable Architecture for RDMA NICsZilong Wang, Layong Luo, Qingsong Ning, Chaoliang Zeng et al.NSDI 2023 · 154 citations
- PSN-PATH: When Multipath RDMA Meets Lossy NetworksZhexiong Li, Shugui Wei, Puyu Zhao, Feifan Wang et al.SIGCOMM 2026
- Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA AtomicGuangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang et al.ISCA 2026
- NetRT: Enhancing RDMA with Retransmission Offloading in Data Center NetworksWentao Wang, Jiangping Han, Kaiping Xue, Jian Li et al.INFOCOM 2025 · 3 citations
- Tlaloc: A Generic Multipath Load Balancing for RoCEHuimin Luo, Jiao Zhang, Yongchen Pan, Tian Pan et al.INFOCOM 2026
