RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency Hiding
Ningzhi Ai, Wenjian He, Hu He, Jing Xia, Heng Liao, Guowei Zhang
摘要
Memory systems characterized by high bandwidth and/or capacity alongside high access latency are becoming increasingly critical.This trend can be observed both at the device level-for instance, in non-volatile memory-and at the system level, as seen in CXL-based memory pooling architectures.To benefit from such memory in general-purpose computing systems, it is essential to employ techniques that can tolerate high memory access latency.Although prefetching has long been recognized as a classical approach for latency tolerance, conventional prefetching techniques are typically either optimized for area efficiency or constrained by limited prefetching patterns.Consequently, they often fail to convert the abundant metadata into significant performance improvements at minimal cost.To address these challenges, we propose RICH-a prefetcher that strategically consumes memory capacity and bandwidth to reduce memory access latency.First, RICH is capable of leveraging abundant metadata to improve performance by integrating spatial prefetching with diverse region sizes and prefetch triggers.Second, RICH implements such metadata with minimal overheads by employing a hierarchical on-chip/off-chip storage mechanism, thereby avoiding both large on-chip storage and critical off-chip accesses.We propose a specific implementation of RICH and evaluate it across a wide range of workloads.With increased memory latency, RICH achieves performance improvements of 8.3% over Bingo and 6.2% over PMP.This highlights the RICH's suitability for future memory systems.In a conventional system, RICH still outperforms Bingo by 3.4%.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Merging Similar Patterns for Hardware PrefetchingShizhi Jiang, Qiusong Yang, Yiwei CiMICRO 2022 · 被引用 27 次
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 被引用 33 次
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAMDusol Lee, Yan Sun, Houxiang Ji, Vinit Gupta 等OSDI 2026
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
