Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal Scheduling
Rishabh R. Iyer, Musa Unal, Marios Kogias, George Candea
摘要
Datacenter applications expect microsecond-scale service times and tightly bound tail latency, with future workloads expected to be even more demanding. To address this challenge, state-of-the-art runtimes employ theoretically optimal scheduling policies, namely a single request queue and strict preemption.
We present Concord, a runtime that demonstrates how forgoing this design-while still closely approximating it-enables a significant improvement in application throughput while maintaining tight tail-latency SLOs. We evaluate Concord on microbenchmarks and Google's LevelDB keyvalue store; compared to the state of the art, Concord improves application throughput by up to 52% on microbenchmarks and by up to 83% on LevelDB, while meeting the same tail-latency SLOs. Unlike the state of the art, Concord is application agnostic and does not rely on the nonstandard use of hardware, which makes it immediately deployable in the public cloud. Concord is publicly available at https://dslab.epfl.ch/research/concord.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Making Kernel Bypass Practical for the Cloud with JunctionJoshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse 等NSDI 2024 · 被引用 57 次
- Making Serverless Pay-For-Use a Reality with LeopardTingjia Cao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Tyler Caraza-HarterNSDI 2025 · 被引用 11 次
- The Benefits and Limitations of User Interrupts for Preemptive Userspace SchedulingLinsong Guo, Danial Zuberi, Tal Garfinkel, Amy OusterhoutNSDI 2025 · 被引用 10 次
- Extended User Interrupts (xUI): Fast and Flexible Notification without PollingBerk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel 等ASPLOS 2025 · 被引用 8 次
- Fast, Flexible, and Practical Kernel ExtensionsKumar Kartikeya Dwivedi, Rishabh R. Iyer, Sanidhya KashyapSOSP 2024 · 被引用 7 次
它引用的顶会 Paper11
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford 等NSDI 2020 · 被引用 96 次
- The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter SystemsIrene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk 等SOSP 2021 · 被引用 83 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao 等MICRO 2021 · 被引用 42 次
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias 等SOSP 2021 · 被引用 39 次
相关 Paper
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin 等HPCA 2024 · 被引用 14 次
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng 等USENIX ATC 2025 · 被引用 1 次
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey 等MICRO 2022 · 被引用 11 次
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu 等OSDI 2020 · 被引用 58 次
- Concord: Rethinking Distributed Coherence for Software Caches in Serverless EnvironmentsJovan Stojkovic, Chloe Alverti, Alan Andrade, Nikoleta Iliakopoulou 等HPCA 2025 · 被引用 5 次
