Efficient Microsecond-scale Blind Scheduling with Tiny Quanta
Zhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro, Amy Ousterhout, Sylvia Ratnasamy, Scott Shenker
Abstract
A longstanding performance challenge in datacenter-based applications is how to efficiently handle incoming client requests that spawn many very short (μs scale) jobs that must be handled with high throughput and low tail latency. When no assumptions are made about the duration of individual jobs, or even about the distribution of their durations, this requires blind scheduling with frequent and efficient preemption, which is not scalably supported for μs-level tasks. We present Tiny Quanta (TQ), a system that enables efficient blind scheduling of μs-level workloads. TQ performs fine-grained preemptive scheduling and does so with high performance via a novel combination of two mechanisms: forced multitasking and two-level scheduling. Evaluations with a wide variety of μs-level workloads show that TQ achieves low tail latency while sustaining 1.2x to 6.8x the throughput of prior blind scheduling systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7e897051-7779-461d-adf8-99e2edad64f2Cited by top-tier papers6
- The Benefits and Limitations of User Interrupts for Preemptive Userspace SchedulingLinsong Guo, Danial Zuberi, Tal Garfinkel, Amy OusterhoutNSDI 2025 · 10 citations
- Harvesting Memory-bound CPU Stall Cycles in Software with MSHZhihong Luo, Sam Son, Sylvia Ratnasamy, Scott ShenkerOSDI 2024 · 5 citations
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng et al.USENIX ATC 2025 · 1 citation
- Rakaia: Scalable In-Kernel Scheduling for TCP-Based RPCsRui Yang, Konstantinos Prasopoulos, Edouard BugnionOSDI 2026
- SBB: Eliminating Centralized Bottlenecks in Userspace Network RuntimeKang Hu, Shuqi Dong, Chuandong Li, Ran Yi et al.OSDI 2026
Related papers
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
- Draconis: Network-Accelerated Scheduling for Microsecond-Scale WorkloadsSreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-KiswanyEuroSys 2024 · 5 citations
- Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant MicroservicesAmirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. WenischHPCA 2020 · 30 citations
- MilliSort and MilliQuery: Large-Scale Data-Intensive Computing in MillisecondsYilong Li, Seo Jin Park, John K. OusterhoutNSDI 2021 · 17 citations
