Svalinn: Overload Control in Large-Scale Servers with Multiple Resource Bottlenecks
Bhaskar Subhash Pardeshi, Peidi Song, Ahmed Saeed
摘要
Modern overload controllers treat application binaries as monoliths and react to aggregate performance, a misconception we call the single-queue fallacy . Real applications have diverse, data-dependent execution paths that stress different resources. Reacting to overall performance forces the controller to focus on the most bottlenecked resource while leaving others underutilized. We present Svalinn, a modular overload controller designed to maximize utilization across multiple potential bottlenecks such as CPU, memory bandwidth, and contended locks. Svalinn separates throughput control and latency control. A credit-based admission controller regulates offered load to maximize a user-defined utility function. Per-bottleneck controllers then enforce latency targets using Active Queue Management (AQM) policies. While AQM is straightforward for resources with explicit software queues, managing memory-bandwidth-intensive operations is challenging due to the absence of such queues. To handle this case, we introduce m s emaphore, which adaptively limits the number of concurrent memory-bandwidth-intensive requests to achieve high memory-bandwidth utilization using the minimum necessary CPU cores. We integrate Svalinn into four applications and two runtimes and show that it improves goodput by up to 6.51× without compromising latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim 等OSDI 2020 · 被引用 97 次
相关 Paper
- Protego: Overload Control for Applications with Unpredictable Lock ContentionInho Cho, Ahmed Saeed, Seo Jin Park, Mohammad Alizadeh 等NSDI 2023 · 被引用 11 次
- Building A CSFQ-Inspired Transport for Switched CXL Memory PoolingZerui Guo, Emily Shriver, Ming LiuNSDI 2026 · 被引用 2 次
- CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency ControlQiaoling Chen, Zhisheng Ye, Tian Tang, Peng Sun 等ICML 2026
- M3: end-to-end memory management in elastic system software stacksDavid Lion, Adrian Chiu, Ding YuanEuroSys 2021 · 被引用 3 次
- Emma: Elastic Multi-Resource Management for Realtime Stream ProcessingRengan Dou, Xin Wang, Richard T. B. MaINFOCOM 2024 · 被引用 2 次
