Fathom: Understanding Datacenter Application Network Performance
Mubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh, Yousuk Seung, Neal Cardwell, Willem de Bruijn, Van Jacobson, Jasleen Kaur, David Wetherall, Amin Vahdat
Abstract
We describe our experience with Fathom, a system for identifying the network performance bottlenecks of any service running in the Google fleet. Fathom passively samples RPCs, the principal unit of work for services. It segments the overall latency into host and network components with kernel and RPC stack instrumentation. It records these detailed latency metrics, along with detailed transport connection state, for every sampled RPC. This lets us determine if the completion is constrained by the client, network or server. To scale while enabling analysis, we also aggregate samples into distributions that retain multi-dimensional breakdowns. This provides us with a macroscopic view of individual services. Fathom runs globally in our datacenters for all production traffic, where it monitors billions of TCP connections 24x7. For five years Fathom has been our primary tool for troubleshooting service network issues and assessing network infrastructure changes. We present case studies to show how it has helped us improve our production services.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9999957a-a765-450f-8cab-d91e145efda6Cited by top-tier papers5
- Minder: Faulty Machine Detection for Large-scale Distributed Model TrainingYangtao Deng, Xiang Shi, Zhuo Jiang, Xingjian Zhang et al.NSDI 2025 · 36 citations
- Evolution of Aegis: Fault Diagnosis for AI Model Training Service in ProductionJianbo Dong, Kun Qian, Pengcheng Zhang, Zhilong Zheng et al.NSDI 2025 · 21 citations
- Preventing Network Bottlenecks: Accelerating Datacenter Services with Hotspot-Aware Placement for Compute and StorageHamid Hajabdolali Bazzaz, Yingjie Bi, Weiwu Pang, Minlan Yu et al.NSDI 2025 · 3 citations
- Learnings from Deploying Network QoS Alignment to Application Priorities for Storage ServicesMatthew Buckley, Parsa Pazhooheshy, Z. Morley Mao, Nandita Dukkipati et al.NSDI 2025 · 3 citations
- CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud GatewaysYunming Xiao, Yinchao Yang, Jiaqi Zheng, Xuqian Li et al.SIGCOMM 2026 · 1 citation
Related papers
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale ServicesSam (Likun) Xi, Alexey Alexandrov, Ali Sheikh, Ning Wang et al.SOSP 2026
- Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Your Microservices in Zero CodeJunxian Shen, Han Zhang, Yang Xiang, Xingang Shi et al.SIGCOMM 2023 · 47 citations
- FBDetect: Catching Tiny Performance Regressions at Hyperscale through In-Production MonitoringDong Young Yoon, Yang Wang, Miao Yu, Elvis Huang et al.SOSP 2024 · 5 citations
