GoPTX: Fine-grained GPU Kernel Fusion by PTX-level Instruction Flow Weaving
Kan Wu, Zejia Lin, Mengyue Xi, Zhongchun Zheng, Wenxuan Pan, Xianwei Zhang, Yutong Lu
Abstract
GPUs have been heavily utilized in diverse applications, and numerous approaches, including kernel fusion, have been proposed to boost GPU efficiency through concurrent kernel execution. However, these approaches generally overlook the opportunities to mitigate warp stalls and improve instruction level parallelism (ILP) in inter-kernel resource sharing. To address this issue, we introduce GOPTX, a novel design for kernel fusion that improves ILP through deliberate weaving instructions at the PTX level. GOPTX establishes a merged control flow graph (CFG) from original kernels, enabling to interleaving of instructions that were sequentially executed by default and minimizing pipeline stalls on data hazards. We further propose a latency-aware instruction weaving algorithm for more efficient instruction scheduling and an adaptive code slicing method to enlarge the scheduling space. Experimental evaluation demonstrates that GOPTX achieves an average speedup of over the baseline concurrent execution, with a maximum improvement of 23%. The hardware resource utilization statistics show significant enhancements in eligible warps per cycle and resource use.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f62f3090-3669-471b-bc61-697a4b96de93Cited by top-tier papers1
Ask how each one uses itRelated papers
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang et al.HPCA 2022 · 42 citations
- Control Flow Divergence Optimization by Exploiting Tensor CoresWeiguang Pang, Xu Jiang, Songran Liu, Lei Qiao et al.DAC 2024 · 1 citation
- SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsShinnung Jeong, Liam Paul Cooper, Ju Min Lee, Heelim Choi et al.HPCA 2025 · 2 citations
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken et al.OSDI 2026 · 9 citations
- Navigator: Dynamic Multi-kernel Scheduling to Improve GPU PerformanceJiho Kim, John Kim, Yongjun ParkDAC 2020 · 9 citations
