Fathom: Understanding Datacenter Application Network Performance
Mubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh, Yousuk Seung, Neal Cardwell, Willem de Bruijn, Van Jacobson, Jasleen Kaur, David Wetherall, Amin Vahdat
摘要
We describe our experience with Fathom, a system for identifying the network performance bottlenecks of any service running in the Google fleet. Fathom passively samples RPCs, the principal unit of work for services. It segments the overall latency into host and network components with kernel and RPC stack instrumentation. It records these detailed latency metrics, along with detailed transport connection state, for every sampled RPC. This lets us determine if the completion is constrained by the client, network or server. To scale while enabling analysis, we also aggregate samples into distributions that retain multi-dimensional breakdowns. This provides us with a macroscopic view of individual services. Fathom runs globally in our datacenters for all production traffic, where it monitors billions of TCP connections 24x7. For five years Fathom has been our primary tool for troubleshooting service network issues and assessing network infrastructure changes. We present case studies to show how it has helped us improve our production services.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang 等NSDI 2025 · 被引用 36 次
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng 等NSDI 2025 · 被引用 21 次
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu 等NSDI 2025 · 被引用 3 次
- Learnings from Deploying Network QoS Alignment to Application Priorities for Storage ServicesMatthew Buckley, Parsa Pazhooheshy, Z. Morley Mao, Nandita Dukkipati 等NSDI 2025 · 被引用 3 次
- CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud GatewaysYunming Xiao, Yinchao Yang, Jiaqi Zheng, Xuqian Li 等SIGCOMM 2026 · 被引用 1 次
相关 Paper
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 等SOSP 2023 · 被引用 31 次
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale ServicesSam (Likun) Xi, Alexey Alexandrov, Ali Sheikh, Ning Wang 等SOSP 2026
- Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero CodeJunxian Shen, Han Zhang, Yang Xiang, Xingang Shi 等SIGCOMM 2023 · 被引用 47 次
- FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production MonitoringDong Young Yoon, Yang Wang, Miao Yu, Elvis Huang 等SOSP 2024 · 被引用 5 次
