CliqueMap: productionizing an RMA-based distributed caching system
Arjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble, Harshad Deshmukh, Dan Gibson, Milo M. K. Martin, Amanda Strominger, Thomas F. Wenisch, Amin Vahdat
Abstract
Distributed in-memory caching is a key component of modern Internet services. Such caches are often accessed via remote procedure call (RPC), as RPC frameworks provide rich support for productionization, including protocol versioning, memory efficiency, autoscaling, and hitless upgrades. However, full-featured RPC limits performance and scalability as it incurs high latencies and CPU overheads. Remote Memory Access (RMA) offers a promising alternative, but meeting productionization requirements can be a significant challenge with RMA-based systems due to limited programmability and narrow RMA primitives.
This paper describes the design, implementation, and experience derived from CliqueMap, a hybrid RMA/RPC caching system. CliqueMap has been in production use in Google's datacenters for over three years, currently serves more than 1PB of DRAM, and underlies several end-user visible services. CliqueMap makes use of performant and efficient RMAs on the critical serving path and judiciously applies RPCs toward other functionality. The design embraces lightweight replication, client-based quoruming, selfvalidating server responses, per-operation client-side retries, and co-design with the network layers. These foci lead to a system resilient to the rigors of production and frequent post-deployment evolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10cd6633-5bc9-4487-8f9c-f2bd264d65caCited by top-tier papers16
- Clio: a hardware-software co-designed disaggregated memory systemZhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang et al.ASPLOS 2022 · 110 citations
- SIEVE is Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web CachesYazhuo Zhang, Juncheng Yang, Yao Yue, Ymir Vigfusson et al.NSDI 2024 · 63 citations
- Aquila: A unified, low-latency fabric for datacenter networksDan Gibson, Hema Hariharan, Eric Lance, Moray McLaren et al.NSDI 2022 · 60 citations
- DINOMO: An Elastic, Scalable, High-Performance Key-Value Store for Disaggregated Persistent MemorySe Kwon Lee, Soujanya Ponnapalli, Sharad Singhal, Marcos K. Aguilera et al.VLDB 2022 · 49 citations
- Carbink: Fault-Tolerant Far MemoryYang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao et al.OSDI 2022 · 35 citations
Builds on3
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- Hermes: A Fast, Fault-Tolerant and Linearizable Replication ProtocolAntonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi et al.ASPLOS 2020 · 47 citations
Related papers
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Serialization/Deserialization-free State Transfer in Serverless WorkflowsFangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen et al.EuroSys 2024 · 31 citations
- Hardware-supported remote persistence for distributed persistent memoryZhuohui Duan, Haodi Lu, Haikun Liu, Xiaofei Liao et al.SC 2021 · 8 citations
- PRISM: Rethinking the RDMA Interface for Distributed SystemsMatthew Burke, Sowmya Dharanipragada, Shannon Joyner, Adriana Szekeres et al.SOSP 2021 · 23 citations
- Birds of a Feather Flock Together: Scaling RDMA RPCs with FlockSumit Kumar Monga, Sanidhya Kashyap, Changwoo MinSOSP 2021 · 34 citations
