GhOST: a GPU Out-of-Order Scheduling Technique for Stall Reduction
Ishita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu, Konstantinos Iliakis, Panagiotis-Eleftherios Eleftherakis, Sotirios Xydis, Dimitrios Soudris, Tyler Sorensen, Simone Campanoni, Tor M. Aamodt, David I. August
摘要
Graphics Processing Units (GPUs) use massive multi-threading coupled with static scheduling to hide instruction latencies. Despite this, memory instructions pose a challenge as their latencies vary throughout the application’s execution, leading to stalls. Out-of-order (OoO) execution has been shown to effectively mitigate these types of stalls. However, prior OoO proposals involve costly techniques such as reordering loads and stores, register renaming, or two-phase execution, amplifying implementation overhead and consequently creating a substantial barrier to adoption in GPUs. This paper introduces GhOST, a minimal yet effective OoO technique for GPUs. Without expensive components, GhOST can manifest a substantial portion of the instruction reorderings found in an idealized OoO GPU. GhOST leverages the decode stage’s existing pool of decoded instructions and the existing issue stage’s information about instructions in the pipeline to select instructions for OoO execution with little additional hardware. A comprehensive evaluation of GhOST and the prior state-of-the-art OoO technique across a range of diverse GPU benchmarks yields two surprising insights: (1) Prior works utilized Nvidia’s intermediate representation PTX for evaluation; however, the optimized static instruction scheduling of the final binary form negates many purported improvements from OoO execution; and (2) The prior state-of-the-art OoO technique results in an average slowdown across this set of benchmarks. In contrast, GhOST achieves a maximum and geometric mean speedup on GPU binaries with only a area increase, surpassing previous techniques without slowing down any of the measured benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo 等ISCA 2022 · 被引用 25 次
- BOW: Breathing Operand Windows to Exploit Bypassing in GPUsHodjat Asghari Esfeden, AmirAli Abdolrashidi, Shafiur Rahman, Daniel Wong 等MICRO 2020 · 被引用 19 次
相关 Paper
- Out-of-order backprop: an effective scheduling technique for deep learningHyungjun Oh, Junyeol Lee, HyeongJu Kim, Jiwon SeoEuroSys 2022 · 被引用 14 次
- sCROOGe: Circuit-level Design and Optimization Framework for RISC-V Out-of-Order GPUsMaria Zerva, Panagiotis-Eleftherios Eleftherakis, Alexis Maras, Konstantinos Iliakis 等ISCA 2026
- Ghost Threading: Helper-Thread Prefetching for Real SystemsYuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick 等MICRO 2025 · 被引用 2 次
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran 等MICRO 2023 · 被引用 11 次
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 被引用 27 次
