EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in Datacenters
Xinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li, Du Liu, Guoyao Xu, Guodong Yang, Liping Zhang, Yuemin Wu, Xiaopeng Yuan, Quan Chen, Minyi Guo
Abstract
The complexity of online applications is rapidly increasing, bringing more sophisticated performance anomalies in today's cloud datacenter. To fully understand application behaviors, we should obtain both inter-service communication data via RPC-level tracing and intra-service execution traces via application-level tracing to precisely reason about event causality. However, the average time overhead of existing intra-service tracing schemes on the traced applications is generally about 5-10%, possibly reaching 18% in the worst case. To realize practical intra-service tracing in shared and stressed datacenters, one must achieve extreme tracing efficiency with an overhead at the per-mille level.
In this work, we present EXIST, an extremely efficient intra-service tracing system based on off-the-shelf hardware tracing capabilities. EXIST consists of three cooperative modules to pursue optimal trade-offs towards extremely low overhead. Firstly, it identifies and eliminates costly tracing control operations to guarantee the performance of the observed applications. Secondly, it allocates limited trace buffer space dynamically based on application status. Thirdly, it
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e56ae21b-c593-493c-9099-577c0f3f8ca6Cited by top-tier papers1
Ask how each one uses itBuilds on21
- kAFL: Hardware-Assisted Feedback Fuzzing for OS KernelsSergej Schumilo, Cornelius Aschermann, Robert Gawlik, Sebastian Schinzel et al.USENIX Security 2017 · 324 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Postmortem Program Analysis with Hardware-Enhanced Post-Crash ArtifactsJun Xu, Dongliang Mu, Xinyu Xing, Peng Liu et al.USENIX Security 2017 · 55 citations
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
Related papers
- QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale ServicesSam (Likun) Xi, Alexey Alexandrov, Ali Sheikh, Ning Wang et al.SOSP 2026
- NRCAC: Non-Intrusive Microservice Root Cause Analysis Framework for Cloud ProvidersYi Zhai, Junzhou Luo, Jianrui LiuINFOCOM 2025 · 3 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Enabling Efficient Mobile Tracing with BTraceJiawei Wang, Nian Liu, Arnau Casadevall-Saiz, Yutao Liu et al.ASPLOS 2025
- ANT-man: towards agile power management in the microservice eraXiaofeng Hou, Chao Li, Jiacheng Liu, Lu Zhang et al.SC 2020 · 35 citations
