Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant Microservices
Amirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. Wenisch
Abstract
Managing tail latency is a primary challenge in designing large-scale Internet services. Queuing is a major contributor to end-to-end tail latency, wherein nominal tasks are enqueued behind rare, long ones, due to Head-of-Line (HoL) blocking. In this paper, we introduce Q-Zilla, a scheduling framework to tackle tail latency from a queuing perspective, and CoreZilla, a microarchitectural instantiation of our framework. On the algorthmic front, we first propose Server-Queue Decoupled Size-Interval Task Assignment (SQD-SITA), an efficient scheduling algorithm to minimize tail latency for high-disparity service distributions. SQD-SITA is inspired by an earlier algorithm, SITA, which explicitly seeks to address HoL blocking by providing an express-lane for short tasks, protecting them from queuing behind rare, long ones. But, SITA requires prior knowledge of task lengths to steer them into their corresponding lane, which is impractical. Furthermore, SITA may underperform an M/G/k system when some lanes become underutilized. In contrast, SQD-SITA uses incremental preemption to avoid the need for a priori task-size information, and dynamically reallocates servers to lanes to increase server utilization with no performance penalty. We then introduce Interruptible SQD-SITA, which further improves tail latency at the cost of additional preemptions. Finally, we describe and evaluate CoreZilla, wherein a multi-threaded core efficiently implements ISQD-SITA in a software-transparent manner at minimal cost. Our evaluation demonstrates that CoreZilla improves tail latency over a conventional SMT core with 2, 4, and 8 contexts by 2.25×, 3.23×, and 4.88×, on average, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 36343e40-404d-4660-9a5d-818df4002160Cited by top-tier papers8
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse et al.SOSP 2021 · 60 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
- Syrup: User-Defined Scheduling Across the StackKostis Kaffes, Jack Tigar Humphries, David Mazières, Christos KozyrakisSOSP 2021 · 35 citations
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen et al.USENIX ATC 2023 · 32 citations
- AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server ApplicationsJawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou et al.MICRO 2022 · 22 citations
Related papers
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
- Efficient Microsecond-scale Blind Scheduling with Tiny QuantaZhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro et al.ASPLOS 2024 · 8 citations
- SKQ: Event Scheduling for Optimizing Tail Latency in a Traditional OS KernelSiyao Zhao, Haoyu Gu, Ali José MashtizadehUSENIX ATC 2021 · 11 citations
- Don't stop me Now: Embedding based Scheduling for LLMSRana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang et al.ICLR 2025
- Load is not what you should balance: Introducing PrequalBartek Wydrowski, Robert Kleinberg, Stephen M. Rumble, Aaron ArcherNSDI 2024 · 22 citations
