LDB: An Efficient Latency Profiling Tool for Multithreaded Applications
Inho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh, Adam Belay
摘要
Maintaining low tail latency is critical for the efficiency and performance of large-scale datacenter systems. Software bugs that cause tail latency problems, however, are notoriously difficult to debug. We present LDB, a new latency profiling tool that aims to overcome this challenge by precisely identifying the specific functions that are responsible for tail latency anomalies. LDB observes the latency of all functions in a running program. It uses a novel, software-only technique called stack sampling, where a busy-spinning stack scanner thread polls lightweight metadata recorded in the call stack, shifting tracing costs away from program threads. In addition, LDB uses event tagging to record requests, inter-thread synchronization, and context switching. This can be used, for example, to generate per-request timelines and to find the root cause of complex tail latency problems such as lock contention in multi-threaded programs. We evaluate LDB with three datacenter applications, finding latency problems in each. Our results further show that LDB produces actionable insights, has low overhead, and can rapidly analyze recordings, making it feasible to use in production settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Neutrino: Fine-grained GPU Kernel Profiling via Programmable ProbingSonglin Huang, Chenshu WuOSDI 2025 · 被引用 5 次
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
它引用的顶会 Paper2
- How to diagnose nanosecond network latencies in rich end-host stacksRoni Haecki, Radhika Niranjan Mysore, Lalith Suresh, Gerd Zellweger 等NSDI 2022 · 被引用 58 次
- Hubble: Performance Debugging with In-Production, Just-In-Time Method Tracing on AndroidYu Luo, Kirk Rodrigues, Cuiqin Li, Feng Zhang 等OSDI 2022 · 被引用 11 次
相关 Paper
- TailTracer: Continuous Tail Tracing for Production UseTianyi Liu, Yi Li, Yiyu Zhang, Zhuangda Wang 等OOPSLA 2025 · 被引用 1 次
- ScalAna: automating scaling loss detection with graph analysisYuyang Jin, Haojie Wang, Teng Yu, Xiongchao Tang 等SC 2020 · 被引用 16 次
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal 等SIGCOMM 2026
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 被引用 11 次
- Using sample-based time series data for automated diagnosis of scalability losses in parallel programsLai Wei, John M. Mellor-CrummeyPPoPP 2020 · 被引用 7 次
