GhOST: a GPU Out-of-Order Scheduling Technique for Stall Reduction
Ishita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu, Konstantinos Iliakis, Panagiotis-Eleftherios Eleftherakis, Sotirios Xydis, Dimitrios Soudris, Tyler Sorensen, Simone Campanoni, Tor M. Aamodt, David I. August
Abstract
Graphics Processing Units (GPUs) use massive multi-threading coupled with static scheduling to hide instruction latencies. Despite this, memory instructions pose a challenge as their latencies vary throughout the application’s execution, leading to stalls. Out-of-order (OoO) execution has been shown to effectively mitigate these types of stalls. However, prior OoO proposals involve costly techniques such as reordering loads and stores, register renaming, or two-phase execution, amplifying implementation overhead and consequently creating a substantial barrier to adoption in GPUs. This paper introduces GhOST, a minimal yet effective OoO technique for GPUs. Without expensive components, GhOST can manifest a substantial portion of the instruction reorderings found in an idealized OoO GPU. GhOST leverages the decode stage’s existing pool of decoded instructions and the existing issue stage’s information about instructions in the pipeline to select instructions for OoO execution with little additional hardware. A comprehensive evaluation of GhOST and the prior state-of-the-art OoO technique across a range of diverse GPU benchmarks yields two surprising insights: (1) Prior works utilized Nvidia’s intermediate representation PTX for evaluation; however, the optimized static instruction scheduling of the final binary form negates many purported improvements from OoO execution; and (2) The prior state-of-the-art OoO technique results in an average slowdown across this set of benchmarks. In contrast, GhOST achieves a maximum and geometric mean speedup on GPU binaries with only a area increase, surpassing previous techniques without slowing down any of the measured benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf129fcb-100c-455c-9d55-b680d7dd7f51Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo et al.ISCA 2022 · 25 citations
- BOW: Breathing Operand Windows to Exploit Bypassing in GPUsHodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong et al.MICRO 2020 · 19 citations
Related papers
- Out-of-order backprop: an effective scheduling technique for deep learningHyungjun Oh, Junyeol Lee, HyeongJu Kim, Jiwon SeoEuroSys 2022 · 14 citations
- sCROOGe: Circuit-level Design and Optimization Framework for RISC-V Out-of-Order GPUsMaria Zerva, Panagiotis-Eleftherios Eleftherakis, Alexis Maras, Konstantinos Iliakis et al.ISCA 2026
- Ghost Threading: Helper-Thread Prefetching for Real SystemsYuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick et al.MICRO 2025 · 2 citations
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 27 citations
