RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency Hiding
Ningzhi Ai, Wenjian He, Hu He, Jing Xia, Heng Liao, Guowei Zhang
Abstract
Memory systems characterized by high bandwidth and/or capacity alongside high access latency are becoming increasingly critical.This trend can be observed both at the device level-for instance, in non-volatile memory-and at the system level, as seen in CXL-based memory pooling architectures.To benefit from such memory in general-purpose computing systems, it is essential to employ techniques that can tolerate high memory access latency.Although prefetching has long been recognized as a classical approach for latency tolerance, conventional prefetching techniques are typically either optimized for area efficiency or constrained by limited prefetching patterns.Consequently, they often fail to convert the abundant metadata into significant performance improvements at minimal cost.To address these challenges, we propose RICH-a prefetcher that strategically consumes memory capacity and bandwidth to reduce memory access latency.First, RICH is capable of leveraging abundant metadata to improve performance by integrating spatial prefetching with diverse region sizes and prefetch triggers.Second, RICH implements such metadata with minimal overheads by employing a hierarchical on-chip/off-chip storage mechanism, thereby avoiding both large on-chip storage and critical off-chip accesses.We propose a specific implementation of RICH and evaluate it across a wide range of workloads.With increased memory latency, RICH achieves performance improvements of 8.3% over Bingo and 6.2% over PMP.This highlights the RICH's suitability for future memory systems.In a conventional system, RICH still outperforms Bingo by 3.4%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Merging Similar Patterns for Hardware PrefetchingShizhi Jiang, Qiusong Yang, Yiwei CiMICRO 2022 · 27 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 97 citations
- MAC: Metadata Acceleration for Sustainable Performance in Big-Data Systems with CXL DRAMDusol Lee, Yan Sun, Houxiang Ji, Vinit Gupta et al.OSDI 2026
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
