RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers
Hang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu, Christos Kozyrakis, Ion Stoica, Xin Jin
Abstract
Low-latency online services have strict Service Level Objectives (SLOs) that require datacenter systems to support high throughput at microsecond-scale tail latency. Dataplane operating systems have been designed to scale up multi-core servers with minimal overhead for such SLOs. However, as application demands continue to increase, scaling up is not enough, and serving larger demands requires these systems to scale out to multiple servers in a rack. We present RackSched, the first rack-level microsecond-scale scheduler that provides the abstraction of a rack-scale computer (i.e., a huge server with hundreds to thousands of cores) to an external service with network-system co-design. The core of RackSched is a two-layer scheduling framework that integrates inter-server scheduling in the top-of-rack (ToR) switch with intra-server scheduling in each server. We use a combination of analytical results and simulations to show that it provides near-optimal performance as centralized scheduling policies, and is robust for both low-dispersion and high-dispersion workloads. We design a custom switch data plane for the inter-server scheduler, which realizes power-of-k-choices, ensures request affinity, and tracks server loads accurately and efficiently. We implement a RackSched prototype on a cluster of commodity servers connected by a Barefoot Tofino switch. End-to-end experiments on a twelve-server testbed show that RackSched improves the throughput by up to 1.44x, and scales out the throughput near linearly, while maintaining the same tail latency as one server until the system is saturated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef27356e-a4ef-449e-a31f-0586a1819f9aCited by top-tier papers24
- Concordia: Distributed Shared Memory with In-Network Cache CoherenceQing Wang, Youyou Lu, Erci Xu, Junru Li et al.FAST 2021 · 74 citations
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse et al.SOSP 2021 · 60 citations
- Efficient Scheduling Policies for Microsecond-Scale TasksSarah McClure, Amy Ousterhout, Scott Shenker, Sylvia RatnasamyNSDI 2022 · 43 citations
- Cilantro: Performance-Aware Resource Allocation for General Objectives via Online FeedbackRomil Bhardwaj, Kirthevasan Kandasamy, Asim Biswal, Wenshuo Guo et al.OSDI 2023 · 41 citations
- Syrup: User-Defined Scheduling Across the StackKostis Kaffes, Jack Tigar Humphries, David Mazières, Christos KozyrakisSOSP 2021 · 35 citations
Builds on2
- NetLock: Fast, Centralized Lock Management Using Programmable SwitchesZhuolong Yu, Yiwen Zhang, Vladimir Braverman, Mosharaf Chowdhury et al.SIGCOMM 2020 · 60 citations
- Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionHang Zhu, Zhihao Bai, Jialin Li, Ellis Michael et al.VLDB 2020 · 58 citations
Related papers
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng et al.USENIX ATC 2025 · 1 citation
- NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCsGyuyeong KimSIGCOMM 2023 · 4 citations
- Horus: Granular In-Network Task Scheduler for Cloud DatacentersParham Yassini, Khaled Diab, Saeed Mahloujifar, Mohamed HefeedaNSDI 2024 · 14 citations
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro et al.NSDI 2023 · 26 citations
- Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal SchedulingRishabh R. Iyer, Musa Unal, Marios Kogias, George CandeaSOSP 2023 · 17 citations
