PD3: Prefetching Data with DPUs for Disaggregated Memory
Sidharth Sankhe, Felix Zhang, Umayrah Chonee, Sherman Lim, Jiasheng Hu, Jialin Li, Qizhen Zhang
Abstract
We introduce PD3, a memory disaggregation solution that avoids cache misses, via prefetching, on compute servers and thus all their associated overhead. Unlike a traditional prefetcher that may pollute the cache or miss preloading opportunities due to false positives and false negatives, PD3 prevents mispredictions with network support and minimal yet critical application information. Enabling PD3 is a data processing unit or DPU, which allows (1) parsing user requests before they are processed by the compute server, (2) fetching data from remote memory on the shortest path, (3) offloading expensive RDMA and DMA operations from the host, and (4) incorporating application knowledge to faithfully predict cache misses and take actions accordingly. Designing PD3 requires reconciling DPU resource constraints and scaling requirements of cloud data systems, as well as achieving high efficiency with a myriad of performance optimizations. Our experimental results on real hardware, applications, and workloads show that with nominal compute-local memory, PD3 minimizes the performance gap between memorydisaggregated applications and their monolithic counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c61394b-4692-4f0d-b2cc-98b45ab9787bCited by top-tier papers1
Ask how each one uses itBuilds on24
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 224 citations
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 186 citations
- One-sided RDMA-Conscious Extendible Hashing for Disaggregated MemoryPengfei Zuo, Jiazhao Sun, Liu Yang, Shuangwu Zhang et al.USENIX ATC 2021 · 113 citations
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im et al.SOSP 2021 · 83 citations
- Managing Memory Tiers with CXL in Virtualized EnvironmentsYuhong Zhong, Daniel S. Berger, Carl A. Waldspurger, Ryan Wee et al.OSDI 2024 · 77 citations
Related papers
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu et al.VLDB 2024 · 12 citations
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 15 citations
- Fast Distributed Transactions for RDMA-based Disaggregated MemoryHaodi Lu, Haikun Liu, Yujian Zhang, Zhuohui Duan et al.USENIX ATC 2025 · 9 citations
- Wiseswap: Elastic Datacenter Network-Aware Disaggregated Memory for Multi-Tenant CloudMingxuan Liu, Jianhua Gu, Tianhai Zhao, Dong SunWWW 2026
- Scalable Distributed Inverted List Indexes in Disaggregated MemoryManuel Widmoser, Daniel Kocher, Nikolaus AugstenSIGMOD 2024 · 5 citations
