APT-GET: profile-guided timely software prefetching
Saba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci, Heiner Litz
Abstract
Prefetching which predicts future memory accesses and preloads them from main memory, is a widely-adopted technique to overcome the processor-memory performance gap. Unfortunately, hardware prefetchers implemented in today's processors cannot identify complex and irregular memory access patterns exhibited by modern data-driven applications and hence developers need to rely on software prefetching techniques. We investigate the challenges of enabling effective, automated software data prefetching. Our investigation reveals that the state-of-the-art compiler-based prefetching mechanism falls short in achieving high performance due to its static nature. Based on this insight, we design APT-GET, a novel profile-guided technique that ensures prefetch timeliness by leveraging dynamic execution time information. APT-GET leverages efficient hardware support such as Intel's Last Branch Record (LBR), for collecting application execution profiles with negligible overhead to characterize the execution time of loads. APT-GET then introduces a novel analytical model to find the optimal prefetch-distance and prefetch injection site based on the collected profile to enable timely prefetches. We study APT-GET in the context of 10 real-world applications and demonstrate that it achieves a speedup of up to 1.98× and of 1.30× on average. By ensuring prefetch timeliness, APT-GET improves the performance by 1.25× over the state-of-the-art software data prefetching mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8f3266f-b480-4ceb-99ac-b2643c2117ccCited by top-tier papers13
- Whisper: Profile-Guided Branch Misprediction Elimination for Data Center ApplicationsTanvir Ahmed Khan, Muhammed Ugur, Krishnendra Nathella, Dam Sunwoo et al.MICRO 2022 · 25 citations
- Thermometer: profile-guided btb replacement for data center applicationsShixin Song, Tanvir Ahmed Khan, Sara Mahdizadeh-Shahri, Akshitha Sriraman et al.ISCA 2022 · 23 citations
- Mira: A Program-Behavior-Guided Far Memory SystemZhiyuan Guo, Zijian He, Yiying ZhangSOSP 2023 · 22 citations
- A New Formulation of Neural Data PrefetchingQuang Duong, Akanksha Jain, Calvin LinISCA 2024 · 16 citations
- OCOLOS: Online COde Layout OptimizationSYuxuan Zhang, Tanvir Ahmed Khan, Gilles Pokam, Baris Kasikci et al.MICRO 2022 · 14 citations
Builds on11
- A hierarchical neural model of data prefetchingZhan Shi, Akanksha Jain, Kevin Swersky, Milad Hashemi et al.ASPLOS 2021 · 100 citations
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 83 citations
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang et al.HPCA 2021 · 62 citations
- I-SPY: Context-Driven Conditional Instruction Prefetching with CoalescingTanvir Ahmed Khan, Akshitha Sriraman, Joseph Devietti, Gilles Pokam et al.MICRO 2020 · 37 citations
- Ripple: Profile-Guided Instruction Cache Replacement for Data Center ApplicationsTanvir Ahmed Khan, Dexin Zhang, Akshitha Sriraman, Joseph Devietti et al.ISCA 2021 · 33 citations
Related papers
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
- RPG2: Robust Profile-Guided Runtime Prefetch GenerationYuxuan Zhang, Nathan Sobotka, Soyoon Park, Saba Jamilan et al.ASPLOS 2024 · 11 citations
- Profile-Guided Temporal PrefetchingMengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang et al.ISCA 2025 · 4 citations
- PF-LLM: Large Language Model Hinted Hardware PrefetchingCeyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai et al.ASPLOS 2026
