Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoS
Han Zhao, Weihao Cui, Quan Chen, Youtao Zhang, Yanchao Lu, Chao Li, Jingwen Leng, Minyi Guo
摘要
The proliferation of machine learning applications has promoted both CUDA Cores and Tensor Cores’ integration to meet their acceleration demands. While studies have shown that co-locating multiple tasks on the same GPU can effectively improve system throughput and resource utilization, existing schemes focus on scheduling the resources of traditional CUDA Cores and thus lack the ability to exploit the parallelism between Tensor Cores and CUDA Cores.In this paper, we propose Tacker, a static kernel fusion and scheduling approach to improve GPU utilization of both types of cores while ensuring the QoS (Quality-of-Service) of co-located tasks. Tacker consists of a Tensor-CUDA Core kernel fuser, a duration predictor for fused kernels, and a runtime QoS-aware kernel manager. The kernel fuser enables the flexible fusion of kernels that use Tensor Cores and CUDA Cores, respectively. The duration predictor precisely predicts the duration of the fused kernels. Finally, the kernel manager invokes the fused kernel or the original kernel based on the QoS headroom of latency-critical tasks to improve the system throughput. Our experimental results show that Tacker improves the throughput of best-effort applications compared with state-of-the-art solutions by 18.6% on average, while ensuring the QoS of latency-critical tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor CoresYuetao Chen, Kun Li, Yuhao Wang, Donglin Bai 等PPoPP 2024 · 被引用 25 次
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim 等USENIX ATC 2024 · 被引用 21 次
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao 等EuroSys 2025 · 被引用 19 次
- M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical TypeWeiming Hu, Haoyan Zhang, Cong Guo, Yu Feng 等HPCA 2025 · 被引用 19 次
- Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning WorkloadsWei Zhao, Anand Jayarajan, Gennady PekhimenkoASPLOS 2025 · 被引用 4 次
相关 Paper
- Compile-Time QoS Scheme for Deep Learning InferencesSungin Hong, Hyunjun Kim, Hwansoo HanSC 2025 · 被引用 1 次
- Deadline-Aware Offloading for High-Throughput AcceleratorsTsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. RogersHPCA 2021 · 被引用 16 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- Navigator: Dynamic Multi-kernel Scheduling to Improve GPU PerformanceJiho Kim, John Kim, Yongjun ParkDAC 2020 · 被引用 9 次
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 被引用 4 次
