SBB: Eliminating Centralized Bottlenecks in Userspace Network Runtime
Kang Hu, Shuqi Dong, Chuandong Li, Ran Yi, Zonghao Zhang, Yiming Yao, Bo An, Jie Zhang, Xiaolin Wang, Yingwei Luo, Zhenlin Wang, Diyu Zhou
Abstract
To achieve high throughput, low latency, and high CPU efficiency, userspace network runtimes must perform three types of scheduling: 1) request preemption to minimize tail latency, 2) CPU allocation among services to avoid wasting CPU cycles when the load is low, and 3) request load balancing across worker cores to achieve work conservation. However, prior designs rely on centralized components, which inevitably become scalability bottlenecks as the number of CPU workers increases, limiting system performance scaling.
This paper presents SBB, a purely decentralized userspace network runtime that simultaneously delivers high performance, high CPU efficiency, and high scalability by advancing in both system mechanism and scheduling policy. For system mechanism, SBB leverages the emerging User Interrupt mechanism in a novel way, delivering two types of device interrupts to the userspace runtime: 1) user-level timer interrupts for request preemption, and 2) user-level NIC interrupts for packet arrival to perform CPU allocation. For scheduling policy, SBB introduces a two-level algorithm, which marries flow migration (for persistent imbalance) with task stealing (for temporary imbalance), challenging the conventional wisdom that centralized load balancing outperforms decentralized approaches. Evaluation shows that SBB achieves 1.7× to 5.2× higher throughput when scaling to 48 cores, while meeting the same tail latency target.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9915a940-d1ea-45bf-a41a-0b5f4721f889Builds on24
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter SystemsIrene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk et al.SOSP 2021 · 83 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Rearchitecting Linux Storage Stack for µs Latency and High ThroughputJaehyun Hwang, Midhul Vuppalapati, Simon Peter, Rachit AgarwalOSDI 2021 · 63 citations
- Making Kernel Bypass Practical for the Cloud with JunctionJoshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse et al.NSDI 2024 · 57 citations
Related papers
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal et al.SIGCOMM 2026
- Skyloft: A General High-Efficient Scheduling Framework in User SpaceYuekai Jia, Kaifu Tian, Yuyang You, Yu Chen et al.SOSP 2024 · 3 citations
- The Benefits and Limitations of User Interrupts for Preemptive Userspace SchedulingLinsong Guo, Danial Zuberi, Tal Garfinkel, Amy OusterhoutNSDI 2025 · 10 citations
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse et al.SOSP 2021 · 60 citations
- Fast Core Scheduling with Userspace Process AbstractionJiazhen Lin, Youmin Chen, Shiwei Gao, Youyou LuSOSP 2024 · 3 citations
