When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with Perséphone
Henri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias, Boon Thau Loo, Linh Thi Xuan Phan, Irene Zhang
摘要
This paper introduces Perséphone, a kernel-bypass OS scheduler designed to minimize tail latency for applications executing at microsecond-scale and exhibiting wide service time distributions. Perséphone integrates a new scheduling policy, Dynamic Application-aware Reserved Cores (DARC), that reserves cores for requests with short processing times. Unlike existing kernel-bypass schedulers, DARC is not work conserving. DARC profiles application requests and leaves a small number of cores idle when no short requests are in the queue, so when short requests do arrive, they are not blocked by longer-running ones. Counter-intuitively, leaving cores idle lets DARC maintain lower tail latencies at higher utilization, reducing the overall number of cores needed to serve the same workloads and consequently better utilizing the datacenter resources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- IceBreaker: warming serverless functions better with heterogeneityRohan Basu Roy, Tirthak Patel, Devesh TiwariASPLOS 2022 · 被引用 151 次
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas 等OSDI 2022 · 被引用 100 次
- Making Kernel Bypass Practical for the Cloud with JunctionJoshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse 等NSDI 2024 · 被引用 57 次
- Efficient Scheduling Policies for Microsecond-Scale TasksSarah McClure, Amy Ousterhout, Scott Shenker, Sylvia RatnasamyNSDI 2022 · 被引用 43 次
- Paella: Low-latency Model Serving with Software-defined GPU SchedulingKelvin K. W. Ng, Henri Maxime Demoulin, Vincent LiuSOSP 2023 · 被引用 30 次
它引用的顶会 Paper4
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant MicroservicesAmirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. WenischHPCA 2020 · 被引用 30 次
- Lightweight Preemptible FunctionsSol Boucher, Anuj Kalia, David G. Andersen, Michael KaminskyUSENIX ATC 2020 · 被引用 20 次
相关 Paper
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng 等USENIX ATC 2025 · 被引用 1 次
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin 等HPCA 2024 · 被引用 14 次
- Harvesting Spare CPU Resources in Container SystemsAdam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran 等NSDI 2026 · 被引用 2 次
- Efficient Microsecond-scale Blind Scheduling with Tiny QuantaZhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro 等ASPLOS 2024 · 被引用 8 次
