HoPP: Hardware-Software Co-Designed Page Prefetching for Disaggregated Memory
Haifeng Li, Ke Liu, Ting Liang, Zuojun Li, Tianyue Lu, Hui Yuan, Yinben Xia, Yungang Bao, Mingyu Chen, Yizhou Shan
摘要
Memory disaggregation is a promising direction to mitigate memory contention in datacenters. To make memory disaggregation practical, prior efforts expose remote memory to applications transparently via virtual memory subsystem's swapping interface. However, due to the semantic gap between OS and applications -OS cannot know the memory accessing sequences of an application but via page faults. This approach has two limitations. First, it learns little from page faults' access history, which leads to sub-optimal prefetching predictions. Second, a page fault can still occur even if there is a prefetch-hit which leads to a large kernel overhead.
To address such limitations, our key insight is to decouple the address capturing from page faults by collecting full memory access traces in the memory controller. Using this idea, we build HoPP -a hardware-software co-designed prefetching framework. HoPP adds hardware modules to the memory controller to feed sufficient hot pages to OS in real-time, which has three benefits in HoPP's software design: 1) it improves existing prefetching algorithms with simple revamps, also offers more insights to build better policies; 2) the prefetch algorithm can run as a separate data path alongside the normal remote data path via page faults, potentially hiding the swap latency from applications, and enabling fine-grained control over prefetching behaviors; 3) the prefetch-hit overhead can be eliminated by early page table entry (PTE) injection, i.e., inject PTE for the prefetched page as soon as it returns. We implemented a proof-of-concept prototype using commodity servers along with a hardwarebased memory tracking tool called HMTT to emulate a modified memory controller. Results show that compared to Fastswap and Leap, HoPP-optimized prefetching algorithm achieves over 90% accuracy and coverage, which leads to up to 59% completion time improvement for various datacenter applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory SystemsYan Sun, Jongyul Kim, Zeduo Yu, Jiyuan Zhang 等ASPLOS 2025 · 被引用 27 次
- FetchBPF: Customizable Prefetching Policies in Linux with eBPFXuechun Cao, Shaurya Patel, Soo-Yee Lim, Xueyuan Han 等USENIX ATC 2024 · 被引用 24 次
- NeoMem: Hardware/Software Co-Design for CXL-Native Memory TieringZhe Zhou, Yiqi Chen, Tao Zhang, Yang Wang 等MICRO 2024 · 被引用 17 次
- SepHash: A Write-Optimized Hash Index On Disaggregated Memory via Separate Segment StructureXinhao Min, Kai Lu, Pengyu Liu, Jiguang Wan 等VLDB 2024 · 被引用 9 次
- Beehive: A Scalable Disaggregated Memory Runtime Exploiting Asynchrony of Multithreaded ProgramsQuanxi Li, Hong Huang, Ying Liu, Yanwen Xia 等NSDI 2025 · 被引用 7 次
它引用的顶会 Paper10
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- Borg: the next generationMuhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque 等EuroSys 2020 · 被引用 323 次
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 被引用 224 次
- Effectively Prefetching Remote Memory with LeapHasan Al Maruf, Mosharaf ChowdhuryUSENIX ATC 2020 · 被引用 186 次
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
相关 Paper
- Rethinking software runtimes for disaggregated memoryIrina Calciu, M. Talha Imran, Ivan Puddu, Sanidhya Kashyap 等ASPLOS 2021 · 被引用 116 次
- PD3: Prefetching Data with DPUs for Disaggregated MemorySidharth Sankhe, Felix Zhang, Umayrah Chonee, Sherman Lim 等NSDI 2026 · 被引用 1 次
- DiLOS: Do Not Trade Compatibility for Performance in Memory DisaggregationWonsup Yoon, Jisu Ok, Jinyoung Oh, Sue Moon 等EuroSys 2023 · 被引用 21 次
- Cowbird: Freeing CPUs to Compute by Offloading the Disaggregation of MemoryXinyi Chen, Liangcheng Yu, Vincent Liu, Qizhen ZhangSIGCOMM 2023 · 被引用 15 次
- Wiseswap: Elastic Datacenter Network-Aware Disaggregated Memory for Multi-Tenant CloudMingxuan Liu, Jianhua Gu, Tianhai Zhao, Dong SunWWW 2026
