ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
Mike Chow, Yang Wang, William Wang, Ayichew Hailu, Rohan Bopardikar, Bin Zhang, Jialiang Qu, David Meisner, Santosh Sonawane, Yunqi Zhang, Rodrigo Paim, Mack Ward
摘要
This paper presents ServiceLab, a large-scale performance testing platform developed at Meta. Currently, the diverse set of applications and ML models it tests consumes millions of machines in production, and each year it detects performance regressions that could otherwise lead to the wastage of millions of machines. A major challenge for ServiceLab is to detect small performance regressions, sometimes as tiny as 0.01%. These minor regressions matter due to our large fleet size and their potential to accumulate over time. For instance, the median regression detected by ServiceLab for our large serverless platform, running on more than half a million machines, is only 0.14%. Another challenge is running performance tests in our private cloud, which, like the public cloud, is a noisy environment that exhibits inherent performance variances even for machines of the same instance type. To address these challenges, we conduct a large-scale study with millions of performance experiments to identify machine factors, such as the kernel, CPU, and datacenter location, that introduce variance to test results. Moreover, we present statistical analysis methods to robustly identify small regressions. Finally, we share our seven years of operational experience in dealing with a diverse set of applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Intent-Driven Network Management with Multi-Agent LLMs: The Confucius FrameworkZhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani 等SIGCOMM 2025 · 被引用 22 次
- FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production MonitoringDong Young Yoon, Yang Wang, Miao Yu, Elvis Huang 等SOSP 2024 · 被引用 5 次
- Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted SketchesAlessandro Cornacchia, Theophilus A. Benson, Muhammad Bilal, Marco CaniniNSDI 2026 · 被引用 3 次
它引用的顶会 Paper8
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflowsDarby Huye, Yuri Shkuro, Raja R. SambasivanUSENIX ATC 2023 · 被引用 65 次
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 等SOSP 2023 · 被引用 31 次
相关 Paper
- Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and ExperiencesNeeraj Kumar, Pol Mauri Ruiz, Vijay Menon, Igor Kabiljo 等OSDI 2024 · 被引用 6 次
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 等NSDI 2025 · 被引用 36 次
- SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersYixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar 等ASPLOS 2026
- LLM4JMH: Studying the Use of LLMs for Generating Java Performance MicrobenchmarksZongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang 等ICSE 2026
- Conveyor: One-Tool-Fits-All Continuous Software Deployment at MetaBoris Grubic, Yang Wang, Tyler Petrochko, Ran Yaniv 等OSDI 2023
