Load is not what you should balance: Introducing Prequal
Bartek Wydrowski, Robert Kleinberg, Stephen M. Rumble, Aaron Archer
Abstract
We present Prequal (Probing to Reduce Queuing and Latency), a load balancer for distributed multi-tenant systems. Prequal aims to minimize real-time request latency in the presence of heterogeneous server capacities and non-uniform, time-varying antagonist load. It actively probes server load to leverage the power of d choices paradigm, extending it with asynchronous and reusable probes. Cutting against received wisdom, Prequal does not balance CPU load, but instead selects servers according to estimated latency and active requests-in-flight (RIF). We explore its major design features on a testbed system and evaluate it on YouTube, where it has been deployed for more than two years. Prequal has dramatically decreased tail latency, error rates, and resource use, enabling YouTube and other production systems at Google to run at much higher utilization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77eb5757-5bbc-4949-8397-7e62c381711dCited by top-tier papers6
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang et al.NSDI 2025 · 12 citations
- High-level Programming for Application NetworksXiangfeng Zhu, Yuyao Wang, Banruo Liu, Yongtong Wu et al.NSDI 2025 · 6 citations
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li et al.FAST 2026 · 1 citation
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
- SLATE: Service Layer Traffic EngineeringGangmuk Lim, Aditya Prerepa, Brighten Godfrey, Radhika MittalNSDI 2026
Builds on3
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu et al.OSDI 2020 · 58 citations
- Balanced Allocations: Caching and Packing, Twinning and ThinningDimitrios Los, Thomas Sauerwald, John SylvesterSODA 2022 · 8 citations
- Balanced Allocations with Heterogeneous Bins: The Power of MemoryDimitrios Los, Thomas Sauerwald, John SylvesterSODA 2023 · 5 citations
Related papers
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
- Capybara: Dynamic Load Balancing with Microsecond-Scale TCP MigrationInho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi et al.SIGCOMM 2026
- Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant MicroservicesAmirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. WenischHPCA 2020 · 30 citations
- Protego: Overload Control for Applications with Unpredictable Lock ContentionInho Cho, Ahmed Saeed, Seo Jin Park, Mohammad Alizadeh et al.NSDI 2023 · 11 citations
- The Fast and The Frugal: Tail Latency Aware Provisioning for Coping with Load VariationsAdithya Kumar, Iyswarya Narayanan, Timothy Zhu, Anand SivasubramaniamWWW 2020 · 20 citations
