BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU Systems
AmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh, Nael B. Abu-Ghazaleh, Daniel Wong
摘要
As modern GPU workloads grow in size and complexity, there is an ever-increasing demand for GPU computational power. Emerging workloads contain hundreds or thousands of GPU kernel launches, which incur high overheads, and exhibit data-dependent behavior between kernels, which requires synchronization, leading to GPU under-utilization. Task-based execution models have been proposed to solve these issues, but they require significant programmer effort to port applications to proprietary task-based programming models in order to specify tasks and task dependencies. To address this need, we propose BlockMaestro, a software-hardware solution that combines command queue reordering, kernel-launch-time static analysis, and runtime hardware support to dynamically identify and resolve thread-block level data dependencies between kernels. Through static analysis of memory access patterns at kernel-launch-time, BlockMaestro can extract inter-kernel thread block-level data dependencies. BlockMaestro also introduces kernel pre-launching to reduce the kernel launch overheads experienced by multiple dependent kernels. Correctness is enforced by dynamically resolving thread block-level data dependency at runtime through hardware support. BlockMaestro achieves an average speedup of 51.76% (up to 2.92x) on data-dependent benchmarks, and requires minimal hardware overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU serversKiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song 等SC 2021 · 被引用 17 次
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao 等HPCA 2026 · 被引用 1 次
相关 Paper
- Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisZhanyuan Di, Leping Wang, En Shao, Zhaojia Ma 等ASPLOS 2025 · 被引用 2 次
- Efficient GPU Multitasking with Morphable KernelsTingxu Ren, Ruwen Fan, Hao Guo, Minhui Xie 等SOSP 2026
- Independent Forward Progress of Work-groupsAlexandru Dutu, Matthew D. Sinclair, Bradford M. Beckmann, David A. Wood 等ISCA 2020 · 被引用 5 次
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody 等ASPLOS 2026 · 被引用 1 次
