Speculative Register Reclamation
Sanyam Mehta
Abstract
Large number of in-flight instructions were envisioned two decades ago. They are finally happening now. While more in-flight instructions enable higher ILP and therefore better single-thread performance, it comes at a price. The price is larger structures within the core such as the physical register file. In this work, we propose to reduce register file size while maintaining (or even increasing) number of in-flight instructions. We leverage the insight that within loops, where most time is spent in general, most logical registers are redefined in the same or immediate next iteration. The physical registers allocated to most of these logical registers can thus be aggressively and speculatively released at redefinition instead of being released when the redefining instruction finally commits as in current designs. For correct mis-speculation recovery, only registers that are actually used (i.e used without prior redefinition) across iterations need to remain allocated beyond redefinition, leading to much reduced register file pressure. We show that using our design, register file sizes can be reduced by 50% while still achieving a 1.05x performance improvement over existing designs on a variety of applications even when other core resources are kept the same. The power consumption among various core structures in reduced by 26% on average. In addition, the performance improvement jumps to 1.14x when this reduction in register file size is complemented with an increase of other structures within the core.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 503a68e3-239e-4a81-95bb-68dab0f8a7c0Cited by top-tier papers1
Ask how each one uses itRelated papers
- Tempranillo: Non-Speculative Early Register ReleaseCarlos Escuin, Paolo Salvatore Galfano, Davide Basilio Bartolini, Leeor Peled et al.HPCA 2026
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
- LoopFrog: In-Core Hint-Based Loop ParallelizationMárton Erdos, Utpal Bora, Akshay Bhosale, Bob Lytton et al.MICRO 2025 · 1 citation
- Multi-Stream Squash Reuse for Control-Independent ProcessorsQingxuan Kang, Trevor E. CarlsonMICRO 2025 · 2 citations
- Leveraging Targeted Value Prediction to Unlock New Hardware Strength Reduction PotentialArthur PeraisMICRO 2021 · 11 citations
