Enabling Branch-Mispredict Level Parallelism by Selectively Flushing Instructions
Stijn Eyerman, Wim Heirman, Sam Van den Steen, Ibrahim Hur
Abstract
Conventionally, branch mispredictions are resolved by flushing wrongly speculated instructions from the reorder buffer and refetching instructions along the correct path. However, a large part of the misspeculated instructions could have reconverged with the correct path and executed correctly. Yet, they are flushed to ensure in-order commit. This inefficiency has been recognized in prior work, which proposes either complex additions to a core to reuse the correctly executed instructions, or less intrusive solutions that only reuse part of the converged instructions.
We propose a hardware-software cooperative mechanism to recover correctly executed instructions, avoiding the need to refetch and re-execute them. It combines relatively limited additions to the core architecture with a high reuse of reconverged instructions. Adding the software hints to enable our mechanism is a similar effort as parallelizing an application, which is already necessary to extract high performance from current multicore processors. We evaluate the technique on emerging graph applications and sorting, applications that are known to perform poorly on conventional CPUs, and report an average 29% increase in performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40b2baca-c526-4638-9759-bd82c51ac9faCited by top-tier papers2
- Extended User Interrupts (xUI): Fast and Flexible Notification without PollingBerk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel et al.ASPLOS 2025 · 8 citations
- Multi-Stream Squash Reuse for Control-Independent ProcessorsQingxuan Kang, Trevor E. CarlsonMICRO 2025 · 2 citations
Builds on4
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 83 citations
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 28 citations
- Auto-Predication of Critical BranchesAdarsh Chauhan, Jayesh Gaur, Zeev Sperber, Franck Sala et al.ISCA 2020 · 8 citations
- NOREBA: a compiler-informed non-speculative out-of-order commit processorAli Hajiabadi, Andreas Diavastos, Trevor E. CarlsonASPLOS 2021 · 7 citations
Related papers
- Timely, Efficient, and Accurate Branch PrecomputationAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattMICRO 2024 · 3 citations
- Speculative Register ReclamationSanyam MehtaHPCA 2023 · 3 citations
- Augmenting the Branch Predictor with a Squashed-Branch Reuse BufferRohit Singh, Jiayang Li, Eric RotenbergISCA 2026
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
- Alternate Path FetchAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattISCA 2024 · 2 citations
