LDB: An Efficient Latency Profiling Tool for Multithreaded Applications
Inho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh, Adam Belay
Abstract
Maintaining low tail latency is critical for the efficiency and performance of large-scale datacenter systems. Software bugs that cause tail latency problems, however, are notoriously difficult to debug. We present LDB, a new latency profiling tool that aims to overcome this challenge by precisely identifying the specific functions that are responsible for tail latency anomalies. LDB observes the latency of all functions in a running program. It uses a novel, software-only technique called stack sampling, where a busy-spinning stack scanner thread polls lightweight metadata recorded in the call stack, shifting tracing costs away from program threads. In addition, LDB uses event tagging to record requests, inter-thread synchronization, and context switching. This can be used, for example, to generate per-request timelines and to find the root cause of complex tail latency problems such as lock contention in multi-threaded programs. We evaluate LDB with three datacenter applications, finding latency problems in each. Our results further show that LDB produces actionable insights, has low overhead, and can rapidly analyze recordings, making it feasible to use in production settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7de4fba0-f90e-4aa9-a045-fb32c271ae76Cited by top-tier papers2
- Neutrino: Fine-grained GPU Kernel Profiling via Programmable ProbingSonglin Huang, Chenshu WuOSDI 2025 · 5 citations
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
Builds on2
- How to diagnose nanosecond network latencies in rich end-host stacksRoni Haecki, Radhika Niranjan Mysore, Lalith Suresh, Gerd Zellweger et al.NSDI 2022 · 58 citations
- Hubble: Performance Debugging with In-Production, Just-In-Time Method Tracing on AndroidYu Luo, Kirk Rodrigues, Cuiqin Li, Feng Zhang et al.OSDI 2022 · 11 citations
Related papers
- TailTracer: Continuous Tail Tracing for Production UseTianyi Liu, Yi Li, Yiyu Zhang, Zhuangda Wang et al.OOPSLA 2025 · 1 citation
- ScalAna: automating scaling loss detection with graph analysisYuyang Jin, Haojie Wang, Teng Yu, Xiongchao Tang et al.SC 2020 · 16 citations
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal et al.SIGCOMM 2026
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 11 citations
- Using sample-based time series data for automated diagnosis of scalability losses in parallel programsLai Wei, John M. Mellor-CrummeyPPoPP 2020 · 7 citations
