BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
AmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh, Nael B. Abu-Ghazaleh, Daniel Wong
Abstract
As modern GPU workloads grow in size and complexity, there is an ever-increasing demand for GPU computational power. Emerging workloads contain hundreds or thousands of GPU kernel launches, which incur high overheads, and exhibit data-dependent behavior between kernels, which requires synchronization, leading to GPU under-utilization. Task-based execution models have been proposed to solve these issues, but they require significant programmer effort to port applications to proprietary task-based programming models in order to specify tasks and task dependencies. To address this need, we propose BlockMaestro, a software-hardware solution that combines command queue reordering, kernel-launch-time static analysis, and runtime hardware support to dynamically identify and resolve thread-block level data dependencies between kernels. Through static analysis of memory access patterns at kernel-launch-time, BlockMaestro can extract inter-kernel thread block-level data dependencies. BlockMaestro also introduces kernel pre-launching to reduce the kernel launch overheads experienced by multiple dependent kernels. Correctness is enforced by dynamically resolving thread block-level data dependency at runtime through hardware support. BlockMaestro achieves an average speedup of 51.76% (up to 2.92x) on data-dependent benchmarks, and requires minimal hardware overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8843dfa5-26a9-4120-aae0-4b4cdadccf17Cited by top-tier papers2
- MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU serversKiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song et al.SC 2021 · 17 citations
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao et al.HPCA 2026 · 1 citation
Related papers
- Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisZhanyuan Di, Leping Wang, En Shao, Zhaojia Ma et al.ASPLOS 2025 · 2 citations
- Efficient GPU Multitasking with Morphable KernelsTingxu Ren, Ruwen Fan, Hao Guo, Minhui Xie et al.SOSP 2026
- Independent Forward Progress of Work-groupsAlexandru Dutu, Matthew D. Sinclair, Bradford M. Beckmann, David A. Wood et al.ISCA 2020 · 5 citations
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
