CliqueMap: productionizing an RMA-based distributed caching system
Arjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble, Harshad Deshmukh, Dan Gibson, Milo M. K. Martin, Amanda Strominger, Thomas F. Wenisch, Amin Vahdat
摘要
Distributed in-memory caching is a key component of modern Internet services. Such caches are often accessed via remote procedure call (RPC), as RPC frameworks provide rich support for productionization, including protocol versioning, memory efficiency, autoscaling, and hitless upgrades. However, full-featured RPC limits performance and scalability as it incurs high latencies and CPU overheads. Remote Memory Access (RMA) offers a promising alternative, but meeting productionization requirements can be a significant challenge with RMA-based systems due to limited programmability and narrow RMA primitives.
This paper describes the design, implementation, and experience derived from CliqueMap, a hybrid RMA/RPC caching system. CliqueMap has been in production use in Google's datacenters for over three years, currently serves more than 1PB of DRAM, and underlies several end-user visible services. CliqueMap makes use of performant and efficient RMAs on the critical serving path and judiciously applies RPCs toward other functionality. The design embraces lightweight replication, client-based quoruming, selfvalidating server responses, per-operation client-side retries, and co-design with the network layers. These foci lead to a system resilient to the rigors of production and frequent post-deployment evolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Clio: a hardware-software co-designed disaggregated memory systemZhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang 等ASPLOS 2022 · 被引用 110 次
- SIEVE is Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web CachesYazhuo Zhang, Juncheng Yang, Yao Yue, Ymir Vigfusson 等NSDI 2024 · 被引用 63 次
- Aquila: A unified, low-latency fabric for datacenter networksDan Gibson, Hema Hariharan, Eric Lance, Moray McLaren 等NSDI 2022 · 被引用 60 次
- DINOMO: An Elastic, Scalable, High-Performance Key-Value Store for Disaggregated Persistent MemorySe Kwon Lee, Soujanya Ponnapalli, Sharad Singhal, Marcos K. Aguilera 等VLDB 2022 · 被引用 49 次
- Carbink: Fault-Tolerant Far MemoryYang Zhou, Hassan M. G. Wassel, Sihang Liu, Jiaqi Gao 等OSDI 2022 · 被引用 35 次
它引用的顶会 Paper3
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- Hermes: A Fast, Fault-Tolerant and Linearizable Replication ProtocolAntonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi 等ASPLOS 2020 · 被引用 47 次
相关 Paper
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 等SOSP 2023 · 被引用 31 次
- Serialization/Deserialization-free State Transfer in Serverless WorkflowsFangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 等EuroSys 2024 · 被引用 31 次
- Hardware-supported remote persistence for distributed persistent memoryZhuohui Duan, Haodi Lu, Haikun Liu, Xiaofei Liao 等SC 2021 · 被引用 8 次
- PRISM: Rethinking the RDMA Interface for Distributed SystemsMatthew Burke, Sowmya Dharanipragada, Shannon Joyner, Adriana Szekeres 等SOSP 2021 · 被引用 23 次
- Birds of a Feather Flock Together: Scaling RDMA RPCs with FlockSumit Kumar Monga, Sanidhya Kashyap, Changwoo MinSOSP 2021 · 被引用 34 次
