Control Flow Divergence Optimization by Exploiting Tensor Cores
Weiguang Pang, Xu Jiang, Songran Liu, Lei Qiao, Kexue Fu, Longxiang Gao, Wang Yi
摘要
Kernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- A Modular Static Cost Analysis for GPU Warp-Level ParallelismGregory Blike, Hannah Zicarelli, Udaya Sathiyamoorthy, Julien Lange 等POPL 2026 · 被引用 1 次
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang 等HPCA 2022 · 被引用 42 次
- Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionLishan Yang, Bin Nie, Adwait Jog, Evgenia SmirniICSE 2021 · 被引用 25 次
- µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUsWenhao Huang, Zhaolin Duan, Laiping Zhao, Yuhao Zhang 等HPCA 2026
- Taming Unstructured Sparsity on GPUs via Latency-Aware OptimizationMaohua Zhu, Yuan XieDAC 2020 · 被引用 3 次
