Independent Forward Progress of Work-groups
Alexandru Dutu, Matthew D. Sinclair, Bradford M. Beckmann, David A. Wood, Marcus Chow
Abstract
GPUs have evolved from providing highly-constrained programmability for a single kernel to using pre-emption to ensure independent forward progress for concurrently executing kernels. However, modern GPUs do not ensure independent forward progress for kernels that use fine-grain synchronization to coordinate inter-work-group execution. Enabling independent forward progress among work-groups (WGs) is challenging as pre-empted kernels may be rescheduled with fewer hardware resources. This can lead to oversubscribed execution scenarios that deadlock current hardware even for correctly written code. Prior work addresses this problem by requiring programmers to specify resource requirements and assuming static resource allocation, which adds scheduling constraints and reduces portability. We propose a family of novel hardware approaches - trading off hardware complexity for performance - that provide independent forward progress in the presence of fine-grain inter-WG synchronization and dynamic resource allocation. Additionally, we propose new waiting atomic instructions compatible with proposed C++ 20 extensions. Our final design, Autonomous Work-Groups (AWG), uses hints from regular and waiting atomics to cooperatively schedule WGs within a kernel, improving efficiency and virtualizing hardware resources. In non-oversubscribed scenarios, AWG outperforms a busy-waiting baseline (which deadlocks in oversubscribed scenarios) by 12× on average for benchmarks that use different mutexes and barriers for fine-grained, WG granularity synchronization. Furthermore, AWG outperforms other solutions that do not deadlock in the oversubscribed case, such as fixed-interval round-robin context switching or naively extending monitor/mwait to GPUs, by 2.6× and 2.2×, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8051ee90-093a-489e-bbc5-71ff9b227fe3Cited by top-tier papers3
- KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference ServersMarcus Chow, Ali Jahanshahi, Daniel WongHPCA 2023 · 25 citations
- MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU serversKiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song et al.SC 2021 · 17 citations
- Specifying and testing GPU workgroup progress modelsTyler Sorensen, Lucas F. Salvador, Harmit Raval, Hugues Evrard et al.OOPSLA 2021 · 11 citations
Related papers
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh et al.ISCA 2021 · 15 citations
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang et al.USENIX ATC 2025 · 1 citation
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen et al.OSDI 2025 · 9 citations
- Deadline-Aware Offloading for High-Throughput AcceleratorsTsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. RogersHPCA 2021 · 16 citations
