Caladan: Mitigating Interference at Microsecond Timescales
Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam Belay
摘要
The conventional wisdom is that CPU resources such as cores, caches, and memory bandwidth must be partitioned to achieve performance isolation between tasks. Both the widespread availability of cache partitioning in modern CPUs and the recommended practice of pinning latency-sensitive applications to dedicated cores attest to this belief.
In this paper, we show that resource partitioning is neither necessary nor sufficient. Many applications experience bursty request patterns or phased behavior, drastically changing the amount and type of resources they need. Unfortunately, partitioning-based systems fail to react quickly enough to keep up with these changes, resulting in extreme spikes in latency and lost opportunities to increase CPU utilization.
Caladan is a new CPU scheduler that can achieve significantly better quality of service (tail latency, throughput, etc.) through a collection of control signals and policies that rely on fast core allocation instead of resource partitioning. Caladan consists of a centralized scheduler core that actively manages resource contention in the memory hierarchy and between hyperthreads, and a kernel module that bypasses the standard Linux Kernel scheduler to support microsecondscale monitoring and placement of tasks. When colocating memcached with a best-effort, garbage-collected workload, Caladan outperforms Parties, a state-of-the-art resource partitioning system, by 11,000×, reducing tail latency from 580 ms to 52 µs during shifts in resource usage while maintaining high CPU utilization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper88
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas 等OSDI 2022 · 被引用 100 次
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- Electrode: Accelerating Distributed Protocols with eBPFYang Zhou, Zezhou Wang, Sowmya Dharanipragada, Minlan YuNSDI 2023 · 被引用 78 次
- Rearchitecting Linux Storage Stack for µs Latency and High ThroughputJaehyun Hwang, Midhul Vuppalapati, Simon Peter, Rachit AgarwalOSDI 2021 · 被引用 63 次
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse 等SOSP 2021 · 被引用 60 次
它引用的顶会 Paper3
- Foreshadow: Extracting the Keys to the Intel SGX Kingdom with Transient Out-of-Order ExecutionJo Van Bulck, Marina Minkin, Ofir Weisse, Daniel Genkin 等USENIX Security 2018 · 被引用 1,175 次
- RIDL: Rogue In-Flight Data LoadStephan van Schaik, Alyssa Milburn, Sebastian Österlund, Pietro Frigo 等S&P 2019 · 被引用 408 次
- Port Contention for Fun and ProfitAlejandro Cabrera Aldaya, Billy Bob Brumley, Sohaib ul Hassan, Cesar Pereida García 等S&P 2019 · 被引用 240 次
相关 Paper
- Efficient Scheduling Policies for Microsecond-Scale TasksSarah McClure, Amy Ousterhout, Scott Shenker, Sylvia RatnasamyNSDI 2022 · 被引用 43 次
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias 等SOSP 2021 · 被引用 39 次
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng 等USENIX ATC 2025 · 被引用 1 次
- Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsLiren Zhu, Liujia Li, Jianyu Wu, Yiming Yao 等HPCA 2025 · 被引用 4 次
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro 等NSDI 2023 · 被引用 26 次
