Lune

DAC2024Top-tier venue

Control Flow Divergence Optimization by Exploiting Tensor Cores

Weiguang Pang, Xu Jiang, Songran Liu, Lei Qiao, Kexue Fu, Longxiang Gao, Wang Yi

2024Year
1Citations

Abstract

Kernels are scheduled on Graphics Processing Units (GPUs) in the granularity of GPU warp, which is a bunch of threads that must be scheduled together. When executing kernels with conditional branches, the threads within a warp may execute different branches sequentially, resulting in a considerable utilization loss and unpredictable execution time. This problem is known as the control flow divergence. In this work, we propose a novel method to predict threads' execution path before the launch of the kernel by deploying a branch prediction network on the GPU's tensor cores, which can efficiently parallel run with the kernels on CUDA cores, so that the divergence problem can be eased in a large extent with the lowest overhead. Combined with a well-designed thread data reorganization algorithm, this solution can better mitigate GPUs' control flow divergence problem.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ecc0145c-19a9-4450-b51d-89857a0b044f

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines