Efficient GPU Multitasking with Morphable Kernels
Tingxu Ren, Ruwen Fan, Hao Guo, Minhui Xie, Shiwei Gao, Jiwu Shu, Youyou Lu
Abstract
GPU multitasking offers a promising approach to improving hardware utilization by co-locating concurrent workloads on the same device. However, achieving high resource utilization with minimized interference requires fine-grained, adaptive scheduling. Existing scheduling solutions are fundamentally constrained by a rigid assumption: once a kernel is launched, its resource footprint remains fixed throughout its execution. Consequently, they either rely on static resource pre-partitioning, which lacks the flexibility to adapt to rapid workload changes, or adopt kernel-slicing techniques, which achieve fine-grained control at the cost of prohibitive kernel launch overheads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3cbba97a-6619-4ad5-bf5a-1316332ab9aaRelated papers
- Navigator: Dynamic Multi-kernel Scheduling to Improve GPU PerformanceJiho Kim, John Kim, Yongjun ParkDAC 2020 · 9 citations
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh et al.ISCA 2021 · 15 citations
- µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUsWenhao Huang, Zhaolin Duan, Laiping Zhao, Yuhao Zhang et al.HPCA 2026
- Interference-aware Multiplexing for Deep Learning in GPU Clusters: A Middleware ApproachWenyan Chen, Zizhao Mo, Huanle Xu, Kejiang Ye et al.SC 2023 · 21 citations
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang et al.HPCA 2022 · 42 citations
