Protego: Overload Control for Applications with Unpredictable Lock Contention
Inho Cho, Ahmed Saeed, Seo Jin Park, Mohammad Alizadeh, Adam Belay
摘要
Modern datacenter applications are concurrent, so they require synchronization to control access to shared data. Requests can contend for different combinations of locks, depending on application and request state. In this paper, we show that locks, especially blocking synchronization, can squander throughput and harm tail latency, even when the CPU is underutilized. Moreover, the presence of a large number of contention points, and the unpredictability in knowing which locks a request will require, make it difficult to prevent contention through overload control using traditional signals such as queueing delay and CPU utilization.
We present Protego, a system that resolves these problems with two key ideas. First, it contributes a new admission control strategy that prevents compute congestion in the presence of lock contention. The key idea is to use marginal improvements in observed throughput, rather than CPU load or latency measurements, within a credit-based admission control algorithm that regulates the rate of incoming requests to a server. Second, it introduces a new latency-aware synchronization abstraction called Active Synchronization Queue Management (ASQM) that allows applications to abort requests if delays exceed latency objectives. We apply Protego to two real-world applications, Lucene and Memcached, and show that it achieves up to 3.3× more goodput and 12.2× lower 99th percentile latency than the state-of-the-art overload control systems while avoiding congestion collapse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang 等NSDI 2025 · 被引用 12 次
- TopFull: An Adaptive Top-Down Overload Control for SLO-Oriented MicroservicesJinwoo Park, Jaehyeong Park, Youngmok Jung, Hwijoon Lim 等SIGCOMM 2024 · 被引用 10 次
- Pushing Performance Isolation Boundaries into Application with pBoxYigong Hu, Gongqi Huang, Peng HuangSOSP 2023 · 被引用 2 次
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng 等USENIX ATC 2025 · 被引用 1 次
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
它引用的顶会 Paper6
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Autopilot: workload autoscaling at GoogleKrzysztof Rzadca, Pawel Findeisen, Jacek Swiderski, Przemyslaw Zych 等EuroSys 2020 · 被引用 299 次
- 1RMA: Re-envisioning Remote Memory Access for Multi-tenant DatacentersArjun Singhvi, Aditya Akella, Dan Gibson, Thomas F. Wenisch 等SIGCOMM 2020 · 被引用 70 次
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park 等OSDI 2020 · 被引用 61 次
相关 Paper
- Mitigating Application Resource Overload with Targeted Task CancellationYigong Hu, Zeyin Zhang, Yicheng Liu, Yile Gu 等SOSP 2025
- Achieving Microsecond-Scale Tail Latency Efficiently with Approximate Optimal SchedulingRishabh R. Iyer, Musa Unal, Marios Kogias, George CandeaSOSP 2023 · 被引用 17 次
- Aequitas: admission control for performance-critical RPCs in datacentersYiwen Zhang, Gautam Kumar, Nandita Dukkipati, Xian Wu 等SIGCOMM 2022 · 被引用 30 次
- FlexGuard: Fast Mutual Exclusion Independent of SubscriptionVictor Laforet, Sanidhya Kashyap, Calin Iorgulescu, Julia Lawall 等SOSP 2025
- Avoiding scheduler subversion using scheduler-cooperative locksYuvraj Patel, Leon Yang, Leo Prasath Arulraj, Andrea C. Arpaci-Dusseau 等EuroSys 2020 · 被引用 6 次
