Horus: Granular In-Network Task Scheduler for Cloud Datacenters
Parham Yassini, Khaled Diab, Saeed Mahloujifar, Mohamed Hefeeda
Abstract
Short-lived tasks are prevalent in modern interactive datacenter applications. However, designing schedulers to assign these tasks to workers distributed across the whole datacenter is challenging, because such schedulers need to make decisions at a microsecond scale, achieve high throughput, and minimize the tail response time. Current task schedulers in the literature are limited to individual racks. We present Horus, a new in-network task scheduler for short tasks that operates at the datacenter scale. Horus efficiently tracks and distributes the worker state among switches, which enables it to schedule tasks in parallel at line rate while optimizing the scheduling quality. We propose a new distributed task scheduling policy that minimizes the state and communication overheads, handles dynamic loads, and does not buffer tasks in switches. We compare Horus against the state-of-the-art in-network scheduler in a testbed with programmable switches as well as using simulations of datacenters with more than 27K hosts and thousands of switches handling diverse and dynamic workloads. Our results show that Horus efficiently scales to large datacenters, and it substantially outperforms the state-of-the-art across all performance metrics, including tail response time and throughput. * For multi-packet tasks, Horus maintains a connection table to ensure task affinity. As discussed in §A.4, applications expected to submit multi-packet tasks set the isLastPacket field to 0 in the Horus header (Figure 3), and only state about such tasks are maintained in the connection table.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Pushing the Limits of In-Network Caching for Key-Value StoresGyuyeong KimNSDI 2025 · 9 citations
- Fast and Scalable In-network Lock Management Using Lock FissionHanze Zhang, Ke Cheng, Rong Chen, Haibo ChenOSDI 2024 · 9 citations
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng et al.USENIX ATC 2025 · 1 citation
Builds on11
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Firecracker: Lightweight Virtualization for Serverless ApplicationsAlexandru Agache, Marc Brooker, Alexandra Iordache, Anthony Liguori et al.NSDI 2020 · 197 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic et al.NSDI 2020 · 100 citations
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel et al.OSDI 2020 · 66 citations
Related papers
- Draconis: Network-Accelerated Scheduling for Microsecond-Scale WorkloadsSreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-KiswanyEuroSys 2024 · 5 citations
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu et al.OSDI 2020 · 58 citations
- Efficient Microsecond-scale Blind Scheduling with Tiny QuantaZhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro et al.ASPLOS 2024 · 8 citations
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse et al.SOSP 2021 · 60 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
