RPG2: Robust Profile-Guided Runtime Prefetch Generation
Yuxuan Zhang, Nathan Sobotka, Soyoon Park, Saba Jamilan, Tanvir Ahmed Khan, Baris Kasikci, Gilles A. Pokam, Heiner Litz, Joseph Devietti
Abstract
Data cache prefetching is a well-established optimization to overcome the limits of the cache hierarchy and keep the processor pipeline fed with data. In principle, accurate, welltimed prefetches can sidestep the majority of cache misses and dramatically improve performance. In practice, however, it is challenging to identify which data to prefetch and when to do so. In particular, data can be easily requested too early, causing eviction of useful data from the cache, or requested too late, failing to avoid cache misses. Competition for limited off-chip memory bandwidth must also be balanced between prefetches and a program's regular "demand" accesses. Due to these challenges, prefetching can both help and hurt performance, and the outcome can depend on program structure, decisions about what to prefetch and when to do it, and, as we demonstrate in a series of experiments, program input, processor microarchitecture, and their interaction as well.
To try to meet these challenges, we have designed the RPG 2 system for online prefetch injection and tuning. RPG 2 is a pure-software system that operates on running C/C++ programs, profiling them, injecting prefetch instructions, and then tuning those prefetches to maximize performance. Across dozens of inputs, we find that RPG 2 can provide speedups of up to 2.15×, comparable to the best profileguided prefetching compilers, but can also respond when
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12b6923d-3add-461b-8635-6b44a421e1f0Cited by top-tier papers6
- Constable: Improving Performance and Power Efficiency by Safely Eliminating Load Instruction ExecutionRahul Bera, Adithya Ranganathan, Joydeep Rakshit, Sujit Mahto et al.ISCA 2024 · 8 citations
- Harvesting Memory-bound CPU Stall Cycles in Software with MSHZhihong Luo, Sam Son, Sylvia Ratnasamy, Scott ShenkerOSDI 2024 · 5 citations
- Profile-Guided Temporal PrefetchingMengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang et al.ISCA 2025 · 4 citations
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et al.ISCA 2026
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai et al.ASPLOS 2026
Builds on9
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 83 citations
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang et al.HPCA 2021 · 62 citations
- I-SPY: Context-Driven Conditional Instruction Prefetching with CoalescingTanvir Ahmed Khan, Akshitha Sriraman, Joseph Devietti, Gilles Pokam et al.MICRO 2020 · 37 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
Related papers
- R-Max: Extending BéLáDy's MIN with Prefetching to Bound Realistic Cache PerformanceLei Wang, Chia-Hang Lee, Maccoy Merrell, Gino Chacon et al.ISCA 2026
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci et al.EuroSys 2022 · 25 citations
- CrossPrefetch: Accelerating I/O Prefetching for Modern StorageShaleen Garg, Jian Zhang, Rekha Pitchumani, Manish Parashar et al.ASPLOS 2024 · 9 citations
- Reducing Load Latency with Cache Level PredictionMajid Jalili, Mattan ErezHPCA 2022 · 17 citations
- Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching EfficiencyMengming Li, Qijun Zhang, Yongqing Ren, Zhiyao XieHPCA 2025 · 6 citations
