Delinquent Loop Pre-execution Using Predicated Helper Threads
Anirudh Seshadri, Eric Rotenberg
Abstract
Branch pre-execution targets delinquent branches that are not predictable by conventional branch predictors. Helper threads attempt to resolve branches ahead of the main thread. Pre-executed branch outcomes are communicated to the main thread’s fetch unit via a global branch queue or local branch queues (one per branch PC). Two key challenges are discussed in this paper. 1) Handling a delinquent branch b2 that is control-dependent on another delinquent branch b1. Prior works that include both branches resort to branch prediction of b1 in the helper thread to determine whether or not to pre-execute b2. But b1 is hard-to-predict and the misprediction bottleneck merely shifts from the main thread to the helper thread. 2) Handling a store instruction that both influences a delinquent branch and is control-dependent on it. We propose predicated helper threads (Phelps) to address these challenges. Phelps constructs a helper thread for each inner loop containing delinquent branches. All delinquent branches, even control-dependent ones (b2), are unconditionally pre-executed in each loop iteration. Per-branch queues are managed in lockstep based on loop iterations, allowing the helper thread to deposit outcomes for both b1 and b2 each iteration and the main thread to consume or ignore b2 outcomes in the correct sequence dictated by b1. The helper thread also retains influential stores for dynamic disambiguation and store-load forwarding. Any such store that is control-dependent on a delinquent branch is predicated on the branch’s outcome, which is necessary because the helper thread no longer has control-flow (except for the loop branch). Phelps also features dual decoupled helper threads for outer-inner loop pairs, for effective branch pre-execution when the inner loop has a short and unpredictable trip count.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get af409060-37b0-441e-b984-efd9f4f8b591Related papers
- Slipstream Processors Revisited: Exploiting Branch SetsVinesh Srinivasan, Rangeen Basu Roy Chowdhury, Eric RotenbergISCA 2020 · 12 citations
- Timely, Efficient, and Accurate Branch PrecomputationAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattMICRO 2024 · 3 citations
- Augmenting the Branch Predictor with a Squashed-Branch Reuse BufferRohit Singh, Jiayang Li, Eric RotenbergISCA 2026
- Branch Runahead: An Alternative to Branch Prediction for Impossible to Predict BranchesStephen Pruett, Yale N. PattMICRO 2021 · 23 citations
- BunnyHop: Exploiting the Instruction PrefetcherZhiyuan Zhang, Mingtian Tao, Sioli O'Connell, Chitchanok Chuengsatiansup et al.USENIX Security 2023
