Augmenting the Branch Predictor with a Squashed-Branch Reuse Buffer
Rohit Singh, Jiayang Li, Eric Rotenberg
摘要
In high-performance superscalar processors, a single mispredicted branch may squash hundreds of instructions after it. Some squashed instructions may be control and data independent of the branch, including younger branches, some of which may have executed prior to the squash. Their squashed outcomes can be used to override the branch predictor when the same dynamic branches are refetched, eliminating mispredictions if the predictor is incorrect. The key challenge lies in associating each dynamic branch on the resolved path with its counterpart on the squashed path, if it exists. Multiple dynamic instances of a branch occur due to loops. A dynamic instance can be uniquely identified by a hierarchical loop iteration identifier, like loopA.iter, A,loopB.iter,B for a branch within a doubly-nested loop. We realize such an identifier compactly using an LFSR-based signature that is augmented as loops are entered and continued, and a small stack that saves and restores the signature before entering and after exiting loops, respectively. The signature plus branch PC identifies the dynamic branch being fetched. A key point is that signatures are invariant, in that the signature of a dynamic branch on the resolved path matches that of its counterpart on the squashed path (if it exists), despite arbitrary differences in control-flow observed on the two paths. We augment a 64KB TAGE-SC-L branch predictor with a Squashed-Branch Reuse Buffer (SBRB) accessed by a hash of signature and branch PC. With 11KB of storage for all components, the SBRB improves performance by 2.08% for the SPEC 2006 and 2017 integer benchmarks (maximum 14.1%), 7.25% for the GAPBS benchmarks (maximum 21.2%), and 4.43% for all benchmarks combined.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- The Last-Level Branch PredictorDavid Schall, Andreas Sandberg, Boris GrotMICRO 2024 · 被引用 10 次
- Multi-Stream Squash Reuse for Control-Independent ProcessorsQingxuan Kang, Trevor E. CarlsonMICRO 2025 · 被引用 2 次
- The Last-Level Branch Predictor RevisitedDavid Schall, Mária Duracková, Boris GrotHPCA 2026
- Delinquent Loop Pre-execution Using Predicated Helper ThreadsAnirudh Seshadri, Eric RotenbergHPCA 2025 · 被引用 1 次
- Speculative Register ReclamationSanyam MehtaHPCA 2023 · 被引用 3 次
