Efficient and Scalable Synchronization via Generalized Cache Coherence
Yanpeng Yu, Seung-seob Lee, Lin Zhong, Anurag Khandelwal
Abstract
We explore the design of efficient and scalable synchronization for disaggregated shared memory. Porting existing synchronization primitives to such architectures results in poor performance scaling due to redundant inter-cache communications, exacerbated by high cache-coherence latency in disaggregated shared memory. Driven by our insight that synchronization is a generalization of cache coherence in time and space, we argue for minimally extending existing cache coherence protocols to support synchronization primitives, thereby eliminating the redundant inter-cache communication inherent in layered synchronization. We propose a novel G eneralized cache- C oherence P rotocol (GCP) that realizes this insight by leveraging wait queues and variable-size cache lines directly at the cache-coherence layer for temporal and spatial generalization, respectively. We have verified GCP’s correctness using model checking. We present Soul, an end-to-end system implementation of GCP atop a disaggregated shared-memory platform. Soul supports popular lock APIs through a user-space library that offers improved performance without requiring any changes to application code. Our evaluation of Soul against state-of-the-art locks shows that it improves the performance of unmodified real-world applications at scale by 1–2 orders of magnitude while incurring <8% storage overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1de25fb-894c-4b7b-9c2c-db71ddffa4ddBuilds on22
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout et al.EuroSys 2020 · 163 citations
- Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value StoresShin-Yeh Tsai, Yizhou Shan, Yiying ZhangUSENIX ATC 2020 · 159 citations
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin et al.OSDI 2020 · 96 citations
Related papers
- Rethinking software runtimes for disaggregated memoryIrina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap et al.ASPLOS 2021 · 116 citations
- Cornus: Atomic Commit for a Cloud DBMS with Storage DisaggregationZhihan Guo, Xinyu Zeng, Kan Wu, Wuh-Chwen Hwang et al.VLDB 2023 · 23 citations
- Cache Coherence Over Disaggregated MemoryRuihong Wang, Jianguo Wang, Walid G. ArefVLDB 2025 · 5 citations
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 9 citations
- Efficient, Scalable, and Fair Locking on Disaggregated Memory with Decentralized CoordinationHanze Zhang, Ke Cheng, Rong Chen, Xingda Wei et al.VLDB 2026
