The Last-Level Branch Predictor
David Schall, Andreas Sandberg, Boris Grot
Abstract
Branch prediction is crucial for modern high-performance processors, ensuring efficient execution by anticipating branch outcomes. Despite decades of research, achieving high prediction accuracy remains challenging, particularly in server workloads, where large branch working sets and hard-to-predict branches are prevalent. State-of-the-art predictors, such as the 64KB TAGE-SC-L design, experience high misprediction rates on server workloads, with 3.6-20% (9.2% on average) of execution cycles wasted due to mispredictions on a modern server CPU. While more predictor capacity can reduce mispredictions by up to 36% in the limit (with infinite storage), realizing meaningful gains in practice requires hundreds of KBs of storage, which is infeasible for a latency- and area-sensitive in-core predictor. This work introduces the Last-Level Branch Predictor (LLBP), a microarchitectural approach that improves branch prediction accuracy through additional high-capacity storage backing the baseline TAGE predictor. LLBP leverages the insight that branches requiring longer histories tend to span multiple program contexts - notionally, function calls. A given program context, which can be thought of as a call chain, localizes the branch prediction state, affording a small number of patterns per context even for hard-to-predict branches. LLBP predicts upcoming contexts and prefetches the associated branch metadata into a small in-core buffer, which is accessed in parallel with the unmodified TAGE predictor. Our results show that a 512KB LLBP backing 64KB TAGE-SC-L reduces MPKI by 0.5-25.9% (avg. 8.9 %) over the baseline without LLBP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baa4a982-c986-4c48-8a77-b4aa2b6cf10aCited by top-tier papers4
- Enabling Ahead Prediction with Practical Energy ConstraintsLingzhe Chester Cai, Aniket Deshmukh, Yale N. PattISCA 2025 · 2 citations
- A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction CachingHenry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati et al.MICRO 2025 · 1 citation
- Drishti: Do Not Forget Slicing While Designing Last-Level Cache Replacement Policies for Many-Core SystemsSweta, Prerna Priyadarshini, Biswabandan PandaMICRO 2025 · 1 citation
- Enhancing Instruction Prefetching via Cache and TLB ManagementAlexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis et al.ISCA 2026
Builds on3
- BranchNet: A Convolutional Neural Network to Predict Hard-To-Predict BranchesSiavash Zangeneh, Stephen Pruett, Sangkug Lym, Yale N. PattMICRO 2020 · 50 citations
- Whisper: Profile-Guided Branch Misprediction Elimination for Data Center ApplicationsTanvir Ahmed Khan, Muhammed Ugur, Krishnendra Nathella, Dam Sunwoo et al.MICRO 2022 · 25 citations
- Warming Up a Cold Front-End with IgniteDavid Schall, Andreas Sandberg, Boris GrotMICRO 2023 · 11 citations
Related papers
- The Last-Level Branch Predictor RevisitedDavid Schall, Mária Duracková, Boris GrotHPCA 2026
- Augmenting the Branch Predictor with a Squashed-Branch Reuse BufferRohit Singh, Jiayang Li, Eric RotenbergISCA 2026
- Branch Runahead: An Alternative to Branch Prediction for Impossible to Predict BranchesStephen Pruett, Yale N. PattMICRO 2021 · 23 citations
- Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server ApplicationsTingji Zhang, Boris Grot, Wenjian He, Yashuai Lv et al.ASPLOS 2025 · 4 citations
- AVM-BTB: Adaptive and Virtualized Multi-level Branch Target BufferYunzhe Liu, Xinyu Li, Tingting Zhang, Tianyi Liu et al.ISCA 2024 · 6 citations
