Precise Runahead Execution
Ajeya Naithani, Josué Feliu, Almutaz Adileh, Lieven Eeckhout
Abstract
Runahead execution improves processor performance by accurately prefetching long-latency memory accesses. When a long-latency load causes the instruction window to fill up and halt the pipeline, the processor enters runahead mode and keeps speculatively executing code to trigger accurate prefetches. A recent improvement tracks the chain of instructions that leads to the long-latency load, stores it in a runahead buffer, and executes only this chain during runahead execution, with the purpose of generating more prefetch requests. Unfortunately, all prior runahead proposals have shortcomings that limit performance and energy efficiency because they release processor state when entering runahead mode and then need to re-fill the pipeline to restart normal operation. Moreover, runahead buffer limits prefetch coverage by tracking only a single chain of instructions that leads to the same long-latency load.
We propose precise runahead execution (PRE) which builds on the key observation that when entering runahead mode, the processor has enough issue queue and physical register file resources to speculatively execute instructions. This mitigates the need to release and re-fill processor state in the ROB, issue queue, and physical register file. In addition, PRE preexecutes only those instructions in runahead mode that lead to full-window stalls, using a novel register renaming mechanism to quickly free physical registers in runahead mode, further improving efficiency and effectiveness. Finally, PRE optionally buffers decoded runahead micro-ops in the front-end to save energy. Our experimental evaluation using a set of memoryintensive applications shows that PRE achieves an additional 18.2% performance improvement over the recent runahead proposals while at the same time reducing energy consumption by 6.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa1912ea-28bc-4762-8bdb-2ef6e8954d8eCited by top-tier papers12
- FOCAL: A First-Order Carbon Model to Assess Processor SustainabilityLieven EeckhoutASPLOS 2024 · 30 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
- Branch Runahead: An Alternative to Branch Prediction for Impossible to Predict BranchesStephen Pruett, Yale N. PattMICRO 2021 · 23 citations
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
- Scalar Vector RunaheadJaime Roelandts, Ajeya Naithani, Sam Ainsworth, Timothy M. Jones et al.MICRO 2024 · 11 citations
Related papers
- Reliability-Aware RunaheadAjeya Naithani, Lieven EeckhoutHPCA 2022 · 2 citations
- Criticality Driven FetchAniket Deshmukh, Yale N. PattMICRO 2021 · 9 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- SPECRUN: The Danger of Speculative Runahead Execution in ProcessorsChaoqun Shen, Gang Qu, Jiliang ZhangDAC 2024 · 1 citation
- Boosting Store Buffer Efficiency with Store-Prefetch BurstsJuan M. Cebrian, Stefanos Kaxiras, Alberto RosMICRO 2020 · 4 citations
