FineMem: Breaking the Allocation Overhead vs. Memory Waste Dilemma in Fine-Grained Disaggregated Memory Management
Xiaoyang Wang, Yongkun Li, Kan Wu, Wenzhe Zhu, Yuqi Li, Yinlong Xu
Abstract
RDMA-enabled memory disaggregation has emerged as an attractive approach to reducing memory costs in modern data centers. While RDMA enables efficient remote read/write operations, it presents challenges in remote memory (de)allocation. Consequently, existing systems adopt coarse-grained allocations (in GBs), leading to memory waste.
We introduce FineMem, an RDMA-connected remote memory management system that enables high-performance, fine-grained memory allocation. FineMem addresses latency and scalability challenges related to fine-grained allocations. It removes RDMA memory region (MR) registration costs from allocation paths through per-compute node MR preregistration, while ensuring remote memory isolation using RDMA memory windows and a trusted allocation service on each compute node. It employs a lock-free, one-sided RDMA-based protocol to allocate memory chunks (e.g., 4KB, 2MB) without involving the memory node's CPU and maintains metadata consistency during compute node failures via logging. We show that FineMem reduces remote memory allocation latency by as much as 95% compared to stateof-the-art remote memory management systems. It enables memory malloc systems, key-value stores systems, and swap systems running on FineMem to achieve low memory waste with minimal overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MoonBright: A GPU Memory Allocator with Device-Side Page Table Materialization and Deferred TLB CoherenceYangyu Zhang, Lei Chen, Chunwei Xia, Shuaijiang Li et al.OSDI 2026
- OneSidedMW: Managing Disaggregated Memory Efficiently, Flexibly, and Securely with RNIC OffloadingZixuan Wang, Jinyu Gu, Xingda Wei, Yubin XiaNSDI 2026
Builds on32
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui et al.FAST 2025 · 337 citations
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner et al.ASPLOS 2023 · 255 citations
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 224 citations
Related papers
- Wiseswap: Elastic Datacenter Network-Aware Disaggregated Memory for Multi-Tenant CloudMingxuan Liu, Jianhua Gu, Tianhai Zhao, Dong SunWWW 2026
- ODRP: On-Demand Remote Paging with Programmable RDMAZixuan Wang, Xingda Wei, Jinyu Gu, Hongrui Xie et al.NSDI 2025 · 7 citations
- Fast Distributed Transactions for RDMA-based Disaggregated MemoryHaodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan et al.USENIX ATC 2025 · 9 citations
- REMON: Remote External Memory Over the NetworkShiquan Zhang, Michail Bachras, Yuqiu Zhang, Yunhao Mao et al.ICDE 2026
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 15 citations
