RingLeader: Efficiently Offloading Intra-Server Orchestration to NICs
Jiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro, Brent E. Stephens, Hassan M. G. Wassel, Aditya Akella
Abstract
Careful orchestration of requests at a datacenter server is crucial to meet tight tail latency requirements and ensure high throughput and optimal CPU utilization. Orchestration is multi-pronged and involves load balancing and scheduling requests belonging to different services across CPU resources, and adapting CPU allocation to request bursts. Centralized intra-server orchestration offers ideal load balancing performance, scheduling precision, and burst-tolerant CPU re-allocation. However, existing software-only approaches fail to achieve ideal orchestration because they have limited scalability and waste CPU resources. We argue for a new approach that offloads intra-server orchestration entirely to the NIC. We present RingLeader, a new programmable NIC with novel hardware units for software-informed request load balancing and programmable scheduling and a new light-weight OS-NIC interface that enables close NIC-CPU coordination and supports NIC-assisted CPU scheduling. Detailed experiments with a 100 Gbps FPGA-based prototype show that we obtain better scalability, efficiency, latency, and throughput than state-of-the-art software-only orchestrators including Shinjuku and Caladan.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05acce92-313d-4a9c-b0ab-1fb48076f045Cited by top-tier papers15
- DINT: Fast In-Kernel Distributed Transactions with eBPFYang Zhou, Xingyu Xiang, Matthew Kiley, Sowmya Dharanipragada et al.NSDI 2024 · 40 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- HiDPU: A DPU-Oriented Hybrid Indexing Scheme for Disaggregated Storage SystemsWenbin Zhu, Zhaoyan Shen, Qian Wei, Renhai Chen et al.FAST 2025 · 11 citations
- Enabling Portable and High-Performance SmartNIC Programs with AlkaliJiaxin Lin, Zhiyuan Guo, Mihir Shah, Tao Ji et al.NSDI 2025 · 11 citations
- MTP: Transport for In-Network ComputingTao Ji, Rohan Vardekar, Balajee Vamanan, Brent E. Stephens et al.NSDI 2025 · 9 citations
Builds on12
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford et al.NSDI 2020 · 96 citations
- The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter SystemsIrene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk et al.SOSP 2021 · 83 citations
- FlexTOE: Flexible TCP Offload with Fine-Grained ParallelismRajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon PeterNSDI 2022 · 66 citations
Related papers
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Turbo: SmartNIC-enabled Dynamic Load Balancing of µs-scale RPCsHamed Seyedroudbari, Srikar Vanavasam, Alexandros DaglisHPCA 2023 · 12 citations
- Fast, Scalable, and Accurate Rate Limiter for RDMA NICsZilong Wang, Xinchen Wan, Luyang Li, Yijun Sun et al.SIGCOMM 2024 · 17 citations
- NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCsGyuyeong KimSIGCOMM 2023 · 4 citations
- Network Load Balancing with In-network Reordering Support for RDMACha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi et al.SIGCOMM 2023 · 110 citations
