A Cloud-Scale Characterization of Remote Procedure Calls
Korakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu, Hassan M. G. Wassel, Soheil Hassas Yeganeh, Alex C. Snoeren, Arvind Krishnamurthy, David E. Culler, Henry M. Levy
Abstract
The global scale and challenging requirements of modern cloud applications have led to the development of complex, widely distributed, service-oriented applications. One enabler of such applications is the remote procedure call (RPC), which provides location-independent communication and hides the myriad of cloud communication complexities and requirements within the RPC stack. Understanding RPCs is thus one key to understanding the behavior of cloud applications. While there have been numerous studies of RPCs in distributed systems, as well as attempts to optimize RPC overheads with both software and hardware, there is still a lack of knowledge about the characteristics of RPCs "in the wild" in the modern cloud environment.
To address this gap, we present, to the best of our knowledge, the first large-scale fleet-wide study of RPCs. Our study is conducted at Google, where we measured the infrastructure supporting Google's user-facing, billion-user web services, such as Google Search, Gmail, Maps, and YouTube, and the information and data management systems that support them. To carry out the study, we examined over 10,000 different RPC methods sampled from over one billion traces, along with statistics collected every 30 minutes over a period of nearly two years. Among other things, we consider the volume, throughput and growth rate of RPCs in the datacenter, the latency of RPCs and their components (the "RPC latency tax"), and the structure of RPC call chains. Our analysis shows that the characteristics, scope and complexity of RPCs at hyperscale differ significantly from the assumptions made
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bfdbd44-5c05-4ef8-a47b-e89446d714c4Cited by top-tier papers21
- eTran: Extensible Kernel Transport with eBPFZhongjie Chen, Qingkai Meng, ChonLam Lao, Yifan Liu et al.NSDI 2025 · 17 citations
- Octopus: Enhancing CXL Memory Pods via Sparse TopologyYuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti, Shuwei Teng et al.NSDI 2026 · 15 citations
- Autobahn: Seamless high speed BFTNeil Giridharan, Florian Suri-Payer, Ittai Abraham, Lorenzo Alvisi et al.SOSP 2024 · 13 citations
- Cloudscape: A Study of Storage Services in Modern Cloud ArchitecturesSambhav Satija, Chenhao Ye, Ranjitha Kosgi, Aditya Jain et al.FAST 2025 · 12 citations
- ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production TestingMike Chow, Yang Wang, William Wang, Ayichew Hailu et al.OSDI 2024 · 11 citations
Builds on16
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford et al.NSDI 2020 · 96 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Lifting the veil on Meta's microservice architecture: Analyses of topology and request workflowsDarby Huye, Yuri Shkuro, Raja R. SambasivanUSENIX ATC 2023 · 65 citations
Related papers
- Fathom: Understanding Datacenter Application Network PerformanceMubashir Adnan Qureshi, Junhua Yan, Yuchung Cheng, Soheil Hassas Yeganeh et al.SIGCOMM 2023 · 7 citations
- QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale ServicesSam (Likun) Xi, Alexey Alexandrov, Ali Sheikh, Ning Wang et al.SOSP 2026
- Ship Compute or Ship Data? Why Not Both?Jie You, Jingfeng Wu, Xin Jin, Mosharaf ChowdhuryNSDI 2021 · 25 citations
- Remote Procedure Call as a Managed System ServiceJingrong Chen, Yongji Wu, Shihan Lin, Yechen Xu et al.NSDI 2023 · 30 citations
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li et al.ASPLOS 2025 · 4 citations
