The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems
Lei Zhang, Zhiqiang Xie, Vaastav Anand, Ymir Vigfusson, Jonathan Mace
摘要
Today's distributed tracing frameworks are ill-equipped to troubleshoot rare edge-case requests. The crux of the problem is a trade-off between specificity and overhead. On the one hand, frameworks can indiscriminately select requests to trace when they enter the system (head sampling), but this is unlikely to capture a relevant edge-case trace because the framework cannot know which requests will be problematic until after-the-fact. On the other hand, frameworks can trace everything and later keep only the interesting edge-case traces (tail sampling), but this has high overheads on the traced application and enormous data ingestion costs.
In this paper we circumvent this trade-off for any edge-case with symptoms that can be programmatically detected, such as high tail latency, errors, and bottlenecked queues. We propose a lightweight and always-on distributed tracing system, Hindsight, which implements a retroactive sampling abstraction: instead of eagerly ingesting and processing traces, Hindsight lazily retrieves trace data only after symptoms of a problem are detected. Hindsight is analogous to a car dash-cam that, upon detecting a sudden jolt in momentum, persists the last hour of footage. Developers using Hindsight receive the exact edge-case traces they desire without undue overhead or dependence on luck. Our evaluation shows that Hindsight scales to millions of requests per second, adds nanosecondlevel overhead to generate trace data, handles GB/s of data per node, transparently integrates with existing distributed tracing systems, and successfully persists full, detailed traces in real-world use cases when edge-case problems are detected.
As demonstration, we apply Hindsight on three use cases corresponding to our running examples. We run experiments on the DeathStar Microservices Benchmark [24], the Hadoop Distributed File System [63], an Alibaba benchmark derived from production traces [42], and on several microbenchmarks. We have integrated Hindsight with OpenTelemetry [52] and as a replacement collection component for X-Trace [23]. Our experimental results show that Hindsight imposes nanosecond-scale overhead when generating trace data, can scale to 55 GB/s of data per node, rapidly reconstructs traces when triggered, and coherently captures problematic traces (>99%), as well as related lateral traces, within 100 ms of identifying a symptom.
In summary, our paper makes the following contributions.
• We describe the retroactive sampling abstraction for capturing traces of symptomatic edge-cases.
• We present the design of Hindsight, a distributed tracing system that implements retroactive sampling. Hindsight is compatible with existing tracing APIs and can be transparently integrated with existing applications.
• We apply Hindsight on real-world use cases and show that efficiently collecting edge-case requests is practical. • We evaluate Hindsight on multiple benchmarks and real systems, showing that it can achieve nanosecond-level overhead on trace data generation and handle GB/s data per node while collecting coherent traces. • We illustrate that Hindsight is compatible and performs better than state-of-the-art tracing systems (X-Trace and Jaeger) with more efficient trace-data generation and lower overhead, while providing edge-case tracing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingYangtao Deng, Lei Zhang, Qinlong Wang, Xiaoyun Zhi 等SOSP 2025 · 被引用 7 次
- Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability AnalysisHaiyu Huang, Cheng Chen, Kunyi Chen, Pengfei Chen 等ASPLOS 2025 · 被引用 6 次
- TraStrainer: Adaptive Sampling for Distributed Traces with System Runtime StateHaiyu Huang, Xiaoyu Zhang, Pengfei Chen, Zilong He 等FSE 2024 · 被引用 4 次
- Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted SketchesAlessandro Cornacchia, Theophilus A. Benson, Muhammad Bilal, Marco CaniniNSDI 2026 · 被引用 3 次
- TORAI: Multi-source Root Cause Analysis for Blind Spots in Microservice Service Call GraphLuan Pham, Huong Ha, Xiuzhen Zhang, Hongyu ZhangFSE 2026 · 被引用 3 次
它引用的顶会 Paper2
- Debugging Transient Faults in Data Centers using Synchronized Network-wide Packet HistoriesPravein Govindan Kannan, Nishant Budhdev, Raj Joshi, Mun Choon ChanNSDI 2021 · 被引用 27 次
- Hubble: Performance Debugging with In-Production, Just-In-Time Method Tracing on AndroidYu Luo, Kirk Rodrigues, Cuiqin Li, Feng Zhang 等OSDI 2022 · 被引用 11 次
相关 Paper
- Tracezip: Efficient Distributed Tracing via Trace CompressionZhuangbin Chen, Junsong Pu, Zibin ZhengISSTA 2025 · 被引用 2 次
- Gleaner: A Semantically-Rich and Efficient Online Sampler for Microservice DiagnosticsYifan Yang, Aoyang Fang, Songhan Zhang, Pinjia HeISSTA 2026
- TracePicker: Optimization-Based Trace Sampling for Microservice-Based SystemsShuaiyu Xie, Jian Wang, Maodong Li, Peiran Chen 等FSE 2025 · 被引用 1 次
- FlowScope: Non-Intrusive Distributed Tracing with Method-Level Delay Estimation for Microservices TroubleshootingYantao Geng, Han Zhang, Zhiheng Wu, Yahui Li 等ICSE 2026
- TraceWeaver: Distributed Request Tracing for Microservices Without Application ModificationSachin Ashok, Vipul Harsh, Brighten Godfrey, Radhika Mittal 等SIGCOMM 2024 · 被引用 14 次
