CASINO Core Microarchitecture: Generating Out-of-Order Schedules Using Cascaded In-Order Scheduling Windows
Ipoom Jeong, Seihoon Park, Changmin Lee, Won Woo Ro
Abstract
The performance gap between in-order (InO) and out-of-order (OoO) cores comes from the ability to dynamically create highly optimized instruction issue schedules. In this work, we observe that a significant amount of performance benefit of OoO scheduling can also be attained by supplementing a traditional InO core with a small and speculative instruction scheduling window, namely SpecInO. SpecInO monitors a small set of instructions ahead of a conventional InO scheduling window, aiming at issuing ready instructions behind long-latency stalls. Simulation results show that SpecInO captures and issues 62% of dynamic instructions out of program order. To this end, we propose a CASINO core microarchitecture that dynamically and speculatively generates OoO schedules with near-InO complexity, using CAScaded IN-Order scheduling windows. A Speculative IQ (S-IQ) issues an instruction if it is ready, or otherwise passes it to the next IQ. At the last IQ, instructions are scheduled in program order along serial dependence chains. The net effect is OoO scheduling via collaboration between cascaded InO IQs. To support speculative execution with minimal cost overhead, we propose a novel register renaming technique that allocates free physical registers only to instructions issued from the S-IQ. The proposed core performs dynamic memory disambiguation via an on-commit value check by extending the store buffer already existing in an InO core. We further optimize energy efficiency by filtering out redundant associative searches performed by speculated loads. In our analysis, CASINO core improves performance by 51% over an InO core (within 10 percentage points of an OoO core), which results in 25% and 42% improvements in energy efficiency over InO and OoO cores, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fb550409-7ecb-4bd4-8bf3-b363b40da1b5Cited by top-tier papers3
- Turnpike: Lightweight Soft Error Resilience for In-Order CoresJianping Zeng, Hongjune Kim, Jaejin Lee, Changhee JungMICRO 2021 · 17 citations
- Clockhands: Rename-free Instruction Set Architecture for Out-of-order ProcessorsToru Koizumi, Ryota Shioya, Shu Sugita, Taichi Amano et al.MICRO 2023 · 4 citations
- Architecting Value Prediction around In-Order ExecutionPierre Ravenel, Arthur Perais, Benoît Dupont de Dinechin, Frédéric PétrotHPCA 2025 · 2 citations
Related papers
- Reconstructing Out-of-Order Issue QueueIpoom Jeong, Jiwon Lee, Myung Kuk Yoon, Won Woo RoMICRO 2022 · 9 citations
- Precise Runahead ExecutionAjeya Naithani, Josué Feliu, Almutaz Adileh, Lieven EeckhoutHPCA 2020 · 32 citations
- ATR: Out-of-Order Register Release Exploiting Atomic RegionsYinyuan Zhao, Surim Oh, Mingsheng Xu, Heiner LitzMICRO 2025 · 2 citations
- Speculative Register ReclamationSanyam MehtaHPCA 2023 · 3 citations
- IDLD: Instantaneous Detection of Leakage and Duplication of Identifiers used for Register RenamingYiannakis Sazeides, Alex Gerber, Ron Gabor, Arkady Bramnik et al.MICRO 2022 · 6 citations
