Svalinn: Overload Control in Large-Scale Servers with Multiple Resource Bottlenecks
Bhaskar Subhash Pardeshi, Peidi Song, Ahmed Saeed
Abstract
Modern overload controllers treat application binaries as monoliths and react to aggregate performance, a misconception we call the single-queue fallacy . Real applications have diverse, data-dependent execution paths that stress different resources. Reacting to overall performance forces the controller to focus on the most bottlenecked resource while leaving others underutilized. We present Svalinn, a modular overload controller designed to maximize utilization across multiple potential bottlenecks such as CPU, memory bandwidth, and contended locks. Svalinn separates throughput control and latency control. A credit-based admission controller regulates offered load to maximize a user-defined utility function. Per-bottleneck controllers then enforce latency targets using Active Queue Management (AQM) policies. While AQM is straightforward for resources with explicit software queues, managing memory-bandwidth-intensive operations is challenging due to the absence of such queues. To handle this case, we introduce m s emaphore, which adaptively limits the number of concurrent memory-bandwidth-intensive requests to achieve high memory-bandwidth utilization using the minimum necessary CPU cores. We integrate Svalinn into four applications and two runtimes and show that it improves goodput by up to 6.51× without compromising latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim et al.OSDI 2020 · 97 citations
Related papers
- Protego: Overload Control for Applications with Unpredictable Lock ContentionInho Cho, Ahmed Saeed, Seo Jin Park, Mohammad Alizadeh et al.NSDI 2023 · 11 citations
- Building A CSFQ-Inspired Transport for Switched CXL Memory PoolingZerui Guo, Emily Shriver, Ming LiuNSDI 2026 · 2 citations
- CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency ControlQiaoling Chen, Zhisheng Ye, Tian Tang, Peng Sun et al.ICML 2026
- M3: end-to-end memory management in elastic system software stacksDavid Lion, Adrian Chiu, Ding YuanEuroSys 2021 · 3 citations
- Emma: Elastic Multi-Resource Management for Realtime Stream ProcessingRengan Dou, Xin Wang, Richard T. B. MaINFOCOM 2024 · 2 citations
