Lune

ISCA2026顶会

R-Max: Extending BéLáDy's MIN with Prefetching to Bound Realistic Cache Performance

Lei Wang, Chia-Hang Lee, Maccoy Merrell, Gino Chacon, Daniel A. Jiménez, Paul V. Gratz

2026年份

摘要

Memory performance continues to lag behind the demand of processing elements, a well-known phenomenon known as the memory wall. Cache prefetching is a well-studied and effective method to bridge this gap. Despite a long history of study and the existence of many prefetchers, an open question remains with respect to the upper bound of performance that might be had from prefetching. A “perfect cache” where all accesses hit is often used as an upper bound. However, as we show, this bound is very unrealistic given bandwidth and miss status holding register (MSHR) constraints. Here, we propose a system, R-Max, to approximate ideal prefetching and replacement policy with realistic constraints on bandwidth, cache structure, and capacity but oracular knowledge of future accesses. We compare R-Max's approximated ideal speedup against the speedup of current state-of- the-art prefetchers to show how much remaining performance gain may be left for prefetching. We show that, for a set of workloads taken from SPEC CPU2017, CVP, GAP and XSBench, up to 299.6% maximum and 72.6% average gains are possible under realistic assumptions for a prefetcher that perfectly predicts future accesses, outperforming current state-of-the-art prefetchers by 60.8%. Interestingly, we see that the workloads where R-Max shows the most potential have little relationship with those where existing prefetchers perform best. Taken together, our results highlight the need for new research into prefetching techniques for these under-exploited workloads.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 47301d69-ad5b-4275-bdae-da10b000c81c

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖