FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production Monitoring
Dong Young Yoon, Yang Wang, Miao Yu, Elvis Huang, Juan Ignacio Jones, Abhinay Kukkadapu, Osman Kocas, Jonathan Wiepert, Kapil Goenka, Sherry Chen, Yanjun Lin, Zhihui Huang
Abstract
This paper presents Meta's FBDetect system, which advances the state of the art in performance regression detection by catching regressions as small as 0.005% in noisy production environments. FBDetect monitors around 800,000 time series covering various types of metrics (e.g., throughput, latency, CPU and memory usage) to detect regressions caused by code or configuration changes in hundreds of services running on millions of servers. FBDetect introduces advanced techniques to capture stack traces fleet-wide, measure fine-grained subroutine-level performance differences, filter out deceptive false-positive regressions, deduplicate correlated regressions, and analyze root causes. Beyond these individual techniques, a key strength of FBDetect over prior work is its battle-tested robustness, proven by seven years of production use, and each year catching regressions that would have wasted millions of servers if left undetected.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ce953c6-b13a-46c5-8e0e-a79b4948da63Cited by top-tier papers3
- SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingWei Liu, Kun Qian, Zhenhua Li, Tianyin Xu et al.SIGCOMM 2025 · 8 citations
- Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted SketchesAlessandro Cornacchia, Theophilus A. Benson, Muhammad Bilal, Marco CaniniNSDI 2026 · 3 citations
- Diagnosing Performance Issues in Application-Defined ResourcesYigong Hu, You-Liang Huang, Haodong Zheng, Yicheng Liu et al.OSDI 2026 · 1 citation
Builds on7
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud InfrastructureZe Li, Qian Cheng, Ken Hsieh, Yingnong Dang et al.NSDI 2020 · 69 citations
- Identifying Software Performance Changes Across Variants and VersionsStefan Mühlbauer, Sven Apel, Norbert SiegmundASE 2020 · 25 citations
- Triangulating Python Performance Issues with SCALENEEmery D. Berger, Sam Stern, Juan Altmayer PizzornoOSDI 2023 · 25 citations
Related papers
- ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production TestingMike Chow, Yang Wang, William Wang, Ayichew Hailu et al.OSDI 2024 · 11 citations
- Early Detection of Performance Regressions by Bridging Local Performance Data and Architectural ModelsLizhi Liao, Simon Eismann, Heng Li, Cor-Paul Bezemer et al.ICSE 2025 · 2 citations
- Fathom: Understanding Datacenter Application Network PerformanceMubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh et al.SIGCOMM 2023 · 7 citations
- IoPV: On Inconsistent Option Performance VariationsJinfu Chen, Zishuo Ding, Yiming Tang, Mohammed Sayagh et al.FSE 2023 · 7 citations
- PerfSig: Extracting Performance Bug Signatures via Multi-modality Causal AnalysisJingzhu He, Yuhang Lin, Xiaohui Gu, Chin-Chia Michael Yeh et al.ICSE 2022 · 9 citations
