GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving
Kan Wu, Zejia Lin, Mengyue Xi, Zhongchun Zheng, Wenxuan Pan, Xianwei Zhang, Yutong Lu
摘要
GPUs have been heavily utilized in diverse applications, and numerous approaches, including kernel fusion, have been proposed to boost GPU efficiency through concurrent kernel execution. However, these approaches generally overlook the opportunities to mitigate warp stalls and improve instruction level parallelism (ILP) in inter-kernel resource sharing. To address this issue, we introduce GOPTX, a novel design for kernel fusion that improves ILP through deliberate weaving instructions at the PTX level. GOPTX establishes a merged control flow graph (CFG) from original kernels, enabling to interleaving of instructions that were sequentially executed by default and minimizing pipeline stalls on data hazards. We further propose a latency-aware instruction weaving algorithm for more efficient instruction scheduling and an adaptive code slicing method to enlarge the scheduling space. Experimental evaluation demonstrates that GOPTX achieves an average speedup of over the baseline concurrent execution, with a maximum improvement of 23%. The hardware resource utilization statistics show significant enhancements in eligible warps per cycle and resource use.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang 等HPCA 2022 · 被引用 42 次
- Control Flow Divergence Optimization by Exploiting Tensor CoresWeiguang Pang, Xu Jiang, Songran Liu, Lei Qiao 等DAC 2024 · 被引用 1 次
- SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsShinnung Jeong, Liam Paul Cooper, Ju Min Lee, Heelim Choi 等HPCA 2025 · 被引用 2 次
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken 等OSDI 2026 · 被引用 9 次
- Navigator: Dynamic Multi-kernel Scheduling to Improve GPU PerformanceJiho Kim, John Kim, Yongjun ParkDAC 2020 · 被引用 9 次
