ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
Mike Chow, Yang Wang, William Wang, Ayichew Hailu, Rohan Bopardikar, Bin Zhang, Jialiang Qu, David Meisner, Santosh Sonawane, Yunqi Zhang, Rodrigo Paim, Mack Ward
Abstract
This paper presents ServiceLab, a large-scale performance testing platform developed at Meta. Currently, the diverse set of applications and ML models it tests consumes millions of machines in production, and each year it detects performance regressions that could otherwise lead to the wastage of millions of machines. A major challenge for ServiceLab is to detect small performance regressions, sometimes as tiny as 0.01%. These minor regressions matter due to our large fleet size and their potential to accumulate over time. For instance, the median regression detected by ServiceLab for our large serverless platform, running on more than half a million machines, is only 0.14%. Another challenge is running performance tests in our private cloud, which, like the public cloud, is a noisy environment that exhibits inherent performance variances even for machines of the same instance type. To address these challenges, we conduct a large-scale study with millions of performance experiments to identify machine factors, such as the kernel, CPU, and datacenter location, that introduce variance to test results. Moreover, we present statistical analysis methods to robustly identify small regressions. Finally, we share our seven years of operational experience in dealing with a diverse set of applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a2d931b-4eeb-4977-9bd7-80e62be027cfCited by top-tier papers3
- Intent-Driven Network Management with Multi-Agent LLMs: The Confucius FrameworkZhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani et al.SIGCOMM 2025 · 22 citations
- FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production MonitoringDong Young Yoon, Yang Wang, Miao Yu, Elvis Huang et al.SOSP 2024 · 5 citations
- Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted SketchesAlessandro Cornacchia, Theophilus A. Benson, Muhammad Bilal, Marco CaniniNSDI 2026 · 3 citations
Builds on8
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflowsDarby Huye, Yuri Shkuro, Raja R. SambasivanUSENIX ATC 2023 · 65 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
Related papers
- Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and ExperiencesNeeraj Kumar, Pol Mauri Ruiz, Vijay Menon, Igor Kabiljo et al.OSDI 2024 · 6 citations
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang et al.NSDI 2025 · 36 citations
- SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersYixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar et al.ASPLOS 2026
- LLM4JMH: Studying the Use of LLMs for Generating Java Performance MicrobenchmarksZongxiong Chen, Derui Zhu, Kundi Yao, Weiyi Shang et al.ICSE 2026
- Conveyor: One-Tool-Fits-All Continuous Software Deployment at MetaBoris Grubic, Yang Wang, Tyler Petrochko, Ran Yaniv et al.OSDI 2023
