CLIP: Load Criticality based Data Prefetching for Bandwidth-constrained Many-core Systems
Biswabandan Panda
摘要
Hardware prefetching is a latency-hiding technique that hides the costly off-chip DRAM accesses. However, stateof-the-art prefetchers fail to deliver performance improvement in the case of many-core systems with constrained DRAM bandwidth. For SPEC CPU2017 homogeneous workloads, the state-of-the-art Berti L1 prefetcher, on a 64-core system with four and eight DRAM channels, incurs performance slowdowns of 24% and 16%, respectively. However, Berti improves performance by 35% if we use an unrealistic configuration of 64 DRAM channels for a 64-core system (one DRAM channel per core).
Prior approaches such as prefetch throttling and critical load prefetching are not effective in the presence of state-of-the-art prefetchers. Existing load criticality predictors fail to detect loads that are critical in the presence of hardware prefetching and the best predictor provides an average critical load prediction accuracy of 41%. Existing prefetch throttling techniques use prefetch accuracy as one of the primary metrics. However, these techniques offer limited benefits for state-ofthe-art prefetchers that deliver high prefetch accuracy and use prefetcher-specific throttling and filtering.
We propose CLIP, a novel load criticality predictor for hardware prefetching with constrained DRAM bandwidth. Our load criticality predictor provides an average accuracy of more than 93% and as high as 100%. CLIP also filters out the critical loads that lead to accurate prefetching. For a 64-core system with eight DRAM channels, CLIP improves the effectiveness of state-ofthe-art Berti prefetcher by 24% and 9% for 45 and 200 64-core homogeneous and heterogeneous workload mixes, respectively. We show that CLIP is equally effective in the presence of other state-of-the-art L1 and L2 prefetchers. Overall, CLIP incurs a storage overhead of 1.56KB/core.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- RPG2: Robust Profile-Guided Runtime Prefetch GenerationYuxuan Zhang, Nathan Sobotka, Soyoon Park, Saba Jamilan 等ASPLOS 2024 · 被引用 11 次
- Scalar Vector RunaheadJaime Roelandts, Ajeya Naithani, Sam Ainsworth, Timothy M. Jones 等MICRO 2024 · 被引用 11 次
- To Cross, or Not to Cross Pages for Prefetching?Georgios Vavouliotis, Martí Torrents, Boris Grot, Kleovoulos Kalaitzidis 等HPCA 2025 · 被引用 10 次
- Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction ExecutionRahul Bera, Adithya Ranganathan, Joydeep Rakshit, Sujit Mahto 等ISCA 2024 · 被引用 8 次
- Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching EfficiencyMengming Li, Qijun Zhang, Yongqing Ren, Zhiyao XieHPCA 2025 · 被引用 6 次
它引用的顶会 Paper4
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Effective Mimicry of Belady's MIN PolicyIshan Shah, Akanksha Jain, Calvin LinHPCA 2022 · 被引用 44 次
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 被引用 33 次
- Register file prefetchingSudhanshu Shukla, Sumeet Bandishte, Jayesh Gaur, Sreenivas SubramoneyISCA 2022 · 被引用 8 次
相关 Paper
- Berti: an Accurate Local-Delta Data PrefetcherAgustín Navarro-Torres, Biswabandan Panda, Jesús Alastruey-Benedé, Pablo Ibáñez 等MICRO 2022 · 被引用 82 次
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi 等MICRO 2021 · 被引用 95 次
- I-POP: Ignite Positive PrefetchersYiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu 等HPCA 2026 · 被引用 1 次
- R-Max: Extending BéLáDy's MIN with Prefetching to Bound Realistic Cache PerformanceLei Wang, Chia-Hang Lee, Maccoy Merrell, Gino Chacon 等ISCA 2026
