Slipstream Processors Revisited: Exploiting Branch Sets
Vinesh Srinivasan, Rangeen Basu Roy Chowdhury, Eric Rotenberg
Abstract
Delinquent branches and loads remain key performance limiters in some applications. One approach to mitigate them is pre-execution. Broadly, there are two classes of pre-execution: one class repeatedly forks small helper threads, each targeting an individual dynamic instance of a delinquent branch or load; the other class begins with two redundant threads in a leader-follower arrangement, and speculatively reduces the leading thread. The objective of this paper is to design a new pre-execution microarchitecture that meets four criteria: (i) retains the simpler coordination of a leader-follower microarchitecture, (ii) is fully automated with just hardware, (iii) targets both branches and loads, (iv) and is effective. We review prior preexecution proposals and show that none of them meet all four criteria. We develop Slipstream 2.0 to meet all four criteria. The key innovation in the space of leader-follower architectures is to remove the forward control-flow slices of delinquent branches and loads, from the leading thread. This innovation overcomes key limitations in the only other hardware-only leader-follower prior works: Slipstream and Dual Core Execution (DCE). Slipstream removes backward slices of confident branches to pre-execute unconfident branches, which is ineffective in phases dominated by unconfident branches when branch pre-execution is most needed. DCE is very effective at tolerating cache-missed loads, unless their dependent branches are mispredicted. Removing forward control-flow slices of delinquent branches and delinquent loads enables two firsts, respectively: (1) leader-follower-style branch pre-execution without relying on confident instruction removal, and (2) tolerance of cache-missed loads that feed mispredicted branches. For SPEC 2006/2017 SimPoints wherein Slipstream 2.0 is auto-enabled, it achieves geomean speedups of 67%, 60%, and 12%, over baseline (one core), Slipstream, and DCE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 497f5e39-a810-4def-aa3b-b9b4b721952dCited by top-tier papers4
- Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCsMarcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao et al.ISCA 2022 · 24 citations
- Branch Runahead: An Alternative to Branch Prediction for Impossible to Predict BranchesStephen Pruett, Yale N. PattMICRO 2021 · 23 citations
- Criticality Driven FetchAniket Deshmukh, Yale N. PattMICRO 2021 · 9 citations
- Timely, Efficient, and Accurate Branch PrecomputationAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattMICRO 2024 · 3 citations
Related papers
- Delinquent Loop Pre-execution Using Predicated Helper ThreadsAnirudh Seshadri, Eric RotenbergHPCA 2025 · 1 citation
- Doppelganger Loads: A Safe, Complexity-Effective Optimization for Secure Speculation SchemesAmund Bergland Kvalsvik, Pavlos Aimoniotis, Stefanos Kaxiras, Magnus SjälanderISCA 2023 · 9 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
