Multi-Stream Squash Reuse for Control-Independent Processors
Qingxuan Kang, Trevor E. Carlson
Abstract
Single-core performance remains crucial for mitigating the serial bottleneck in applications, according to Amdahl's Law. However, hard-to-predict branches pose significant challenges to achieve high Instruction-Level Parallelism (ILP) due to frequent pipeline flushes. In typical processors, when a branch is mispredicted, subsequent instructions are indiscriminately flushed, including potentially useful instructions that will be invariably executed in the future, known as Control-Independent (CI) instructions. While existing CI proposals are effective at leveraging idiomatic control flow structures, these techniques overlook more general reconvergence scenarios. Our analysis reveals that due to the program's dynamic execution behavior, the redirected instruction stream can reconverge with not only the last squashed stream, but also several preceding streams. On average, 10% (and up to 31%) of reconvergence opportunities are overlooked if we only consider the interactions between the last squashed instruction stream and the current one.
In this paper, we introduce Multi-Stream Squash Reuse, which discovers execution reuse opportunities from several previously squashed streams. We use tagged rename mapping to enable pairwise comparisons between any two program execution states to identify data reuse. Our proposal achieves average improvements in IPC of 2.2% (SPECint2006), 0.8% (SPECint2017) and 2.4% (GAP) and maximum gains of 8.9% (astar), 6.1% (bc) and 4.0% (cc) from SPECint2006 and GAP benchmark suites.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d5a1cc1-96d1-4c60-b3d4-fe61def2b915Builds on2
- Towards Developing High Performance RISC-V Processors Using Agile MethodologyYinan Xu, Zihao Yu, Dan Tang, Guokai Chen et al.MICRO 2022 · 108 citations
- Enabling Branch-Mispredict Level Parallelism by Selectively Flushing InstructionsStijn Eyerman, Wim Heirman, Sam Van den Steen, Ibrahim HurMICRO 2021 · 10 citations
Related papers
- Augmenting the Branch Predictor with a Squashed-Branch Reuse BufferRohit Singh, Jiayang Li, Eric RotenbergISCA 2026
- Speculative Register ReclamationSanyam MehtaHPCA 2023 · 3 citations
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 28 citations
- Slipstream Processors Revisited: Exploiting Branch SetsVinesh Srinivasan, Rangeen Basu Roy Chowdhury, Eric RotenbergISCA 2020 · 12 citations
- ATR: Out-of-Order Register Release Exploiting Atomic RegionsYinyuan Zhao, Surim Oh, Mingsheng Xu, Heiner LitzMICRO 2025 · 2 citations
