Load is not what you should balance: Introducing Prequal
Bartek Wydrowski, Robert Kleinberg, Stephen M. Rumble, Aaron Archer
摘要
We present Prequal (Probing to Reduce Queuing and Latency), a load balancer for distributed multi-tenant systems. Prequal aims to minimize real-time request latency in the presence of heterogeneous server capacities and non-uniform, time-varying antagonist load. It actively probes server load to leverage the power of d choices paradigm, extending it with asynchronous and reusable probes. Cutting against received wisdom, Prequal does not balance CPU load, but instead selects servers according to estimated latency and active requests-in-flight (RIF). We explore its major design features on a testbed system and evaluate it on YouTube, where it has been deployed for more than two years. Prequal has dramatically decreased tail latency, error rates, and resource use, enabling YouTube and other production systems at Google to run at much higher utilization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Rajomon: Decentralized and Coordinated Overload Control for Latency-Sensitive MicroservicesJiali Xing, Akis Giannoukos, Paul Loh, Shuyue Wang 等NSDI 2025 · 被引用 12 次
- High-level Programming for Application NetworksXiangfeng Zhu, Yuyao Wang, Banruo Liu, Yongtong Wu 等NSDI 2025 · 被引用 6 次
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li 等FAST 2026 · 被引用 1 次
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
- SLATE: Service Layer Traffic EngineeringGangmuk Lim, Aditya Prerepa, Brighten Godfrey, Radhika MittalNSDI 2026
它引用的顶会 Paper3
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu 等OSDI 2020 · 被引用 58 次
- Balanced Allocations: Caching and Packing, Twinning and ThinningDimitrios Los, Thomas Sauerwald, John SylvesterSODA 2022 · 被引用 8 次
- Balanced Allocations with Heterogeneous Bins: The Power of MemoryDimitrios Los, Thomas Sauerwald, John SylvesterSODA 2023 · 被引用 5 次
相关 Paper
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin 等HPCA 2024 · 被引用 14 次
- Capybara: Dynamic Load Balancing with Microsecond-Scale TCP MigrationInho Choi, Nimish Wadekar, Guangda Sun, Raj Joshi 等SIGCOMM 2026
- Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant MicroservicesAmirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. WenischHPCA 2020 · 被引用 30 次
- Protego: Overload Control for Applications with Unpredictable Lock ContentionInho Cho, Ahmed Saeed, Seo Jin Park, Mohammad Alizadeh 等NSDI 2023 · 被引用 11 次
- The Fast and The Frugal: Tail Latency Aware Provisioning for Coping with Load VariationsAdithya Kumar, Iyswarya Narayanan, Timothy Zhu, Anand SivasubramaniamWWW 2020 · 被引用 20 次
