Exploring Instruction Fusion Opportunities in General Purpose Processors
Sawan Singh, Arthur Perais, Alexandra Jimborean, Alberto Ros
Abstract
The Complex Instruction Set Computer (CISC) paradigm has led to the introduction of instruction cracking in which an architectural instruction is divided into multiple microarchitectural instructions (-ops). However, the dual concept, instruction fusion is also prevalent in modern microarchitectures to maximize resource utilization. In essence, some architectural instructions are too complex to be executed as a unit, so they should be cracked, while others are too simple to waste resources on executing them as a unit, so they should be fused with others. In this paper, we focus on instruction fusion and explore opportunities for fusing additional instructions in a high-performance general purpose pipeline. We show that enabling fusion for common RISC-V idioms improves performance by 7%. Then, we determine experimentally that enabling fusion only for memory instructions achieves 86% of the potential of fusion in this particular case. Finally, we propose the Helios microarchitecture, able to fuse non-consecutive and noncontiguous memory instructions, and discuss microarchitectural changes required to do so efficiently while preserving correctness. Helios allows to fuse an additional 5.5% of dynamic instructions, yielding a 14.2% performance uplift over no fusion (8.2% over baseline fusion).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a2b9088-e451-4c3b-8217-748bd7344509Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Improving the Utilization of Micro-operation Caches in x86 ProcessorsJagadish B. Kotra, John KalamatianosMICRO 2020 · 8 citations
- Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V CoresLuca Colagrande, Luca BeniniDAC 2025 · 2 citations
- ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsVincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone et al.DAC 2025 · 1 citation
- GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow WeavingKan Wu, Zejia Lin, Mengyue Xi, Zhongchun Zheng et al.DAC 2025 · 1 citation
- Co-Utilizing SIMD and Scalar to Accelerate the Data Analytics WorkloadsZewen Sun, Zhifang Li, Chuliang WengICDE 2023
