Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware Prefetching
Samuel Pakalapati, Biswabandan Panda
Abstract
Hardware prefetching is one of the common off-chip DRAM latency hiding techniques. Though hardware prefetchers are ubiquitous in the commercial machines and prefetching techniques are well studied in the computer architecture community, the "memory wall" problem still exists after decades of microarchitecture research and is considered to be an essential problem to solve. In this paper, we make a case for breaking the memory wall through data prefetching at the L1 cache.
We propose a bouquet of hardware prefetchers that can handle a variety of access patterns driven by the control flow of an application. We name our proposal Instruction Pointer Classifier based spatial Prefetching (IPCP). We propose IPCP in two flavors: (i) an L1 spatial data prefetcher that classifies instruction pointers at the L1 cache level, and issues prefetch requests based on the classification, and (ii) a multi-level IPCP where the IPCP at the L1 communicates the classification information to the L2 IPCP so that it can kick-start prefetching based on this classification done at the L1. Overall, IPCP is a simple, lightweight, and modular framework for L1 and multi-level spatial prefetching. IPCP at the L1 and L2 incurs a storage overhead of 740 bytes and 155 bytes, respectively.
Our empirical results show that, for memory-intensive singlethreaded SPEC CPU 2017 benchmarks, compared to a baseline system with no prefetching, IPCP provides an average performance improvement of 45.1%. For the entire SPEC CPU 2017 suite, it provides an improvement of 22%. In the case of multicore systems, IPCP provides an improvement of 23.4% (evaluated over more than 1000 mixes). IPCP outperforms the already highperforming state-of-the-art prefetchers like SPP with PPF and Bingo by demanding 30X to 50X less storage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76374806-fea7-48cd-9621-b33973021e5dCited by top-tier papers28
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi et al.MICRO 2021 · 95 citations
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez et al.MICRO 2022 · 82 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
- Page Size Aware Cache PrefetchingGeorgios Vavouliotis, Gino Chacon, Lluc Alvarez, Paul V. Gratz et al.MICRO 2022 · 24 citations
Related papers
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et al.ISCA 2026
- STEP: Spatial Footprint Prefetcher with Multi-Point Temporal TriggersYuanji Ye, Oliver Lenke, Thomas Wild, Andreas HerkersdorfISCA 2026
- CLIP: Load Criticality based Data Prefetching for Bandwidth-constrained Many-core SystemsBiswabandan PandaMICRO 2023 · 21 citations
- Merging Similar Patterns for Hardware PrefetchingShizhi Jiang, Qiusong Yang, Yiwei CiMICRO 2022 · 27 citations
- R-Max: Extending BéLáDy's MIN with Prefetching to Bound Realistic Cache PerformanceLei Wang, Chia-Hang Lee, Maccoy Merrell, Gino Chacon et al.ISCA 2026
