FORGE: Mitigating Synchronization Amplification for Memory-Disaggregated Caching Systems
Zhijun Yang, Yu Hua, Ming Zhang, Menglei Chen, Yixiao Wang
Abstract
Disaggregated Memory (DM) architectures offer caching systems the potential for elastic scaling and improved resource utilization by decoupling compute and memory. However, this advantage is undermined by costly cross-node synchronization, which exacerbates the overheads of critical cache operations, including hotness tracking, eviction coordination, and memory defragmentation. To address this challenge, we present FORGE, a caching system tailored for DM that prioritizes synchronization efficiency. FORGE groups cached objects based on similarity and performs group-level synchronizations to amortize overheads. It evicts cold groups via a contention-free and hotness-aware FIFO queue, efficiently sustaining high hit ratios while mitigating memory fragmentation. Leveraging the predictability of FIFO evictions, FORGE adopts a lazy synchronization strategy that updates hotness metrics just-in-time for eviction and offloads this process to on-chip memory in RDMA NICs for acceleration. Extensive evaluations on YCSB and real-world workloads demonstrate that FORGE achieves up to 4.5× higher throughput, 4.0×/7.5× lower P50/P99 latency, and an average of 1.14× higher cache hit ratio compared with state-of-the-art systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cebfa9ce-23eb-4e0f-9581-eb313fefe0c6Builds on36
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui et al.FAST 2025 · 337 citations
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
Related papers
- Shard: A Scalable and Resize-optimized Hash Index on Disaggregated MemoryHantian Zha, Teng Ma, Baotong Lu, Yuansen Wang et al.VLDB 2026
- DMTree: Towards Efficient Tree Indexing on Disaggregated Memory via Compute-side Collaborative DesignGuoli Wei, Yongkun Li, Haoze Song, Tao Li et al.FAST 2026 · 1 citation
- Fast Distributed Transactions for RDMA-based Disaggregated MemoryHaodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan et al.USENIX ATC 2025 · 9 citations
- CoRM: Compactable Remote Memory over RDMAKonstantin Taranov, Salvatore Di Girolamo, Torsten HoeflerSIGMOD 2021 · 18 citations
- UniMem: Redesigning Disaggregated Memory within A Unified Local-Remote Memory HierarchyYijie Zhong, Minqiang Zhou, Zhirong Shen, Jiwu ShuUSENIX ATC 2024 · 7 citations
