Lune

SC2024Top-tier venue

CUDASTF: Bridging the Gap Between CUDA and Task Parallelism

Cédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael Garland

2024Year
7Citations
1Top-tier citations

Abstract

Organizing computation as asynchronous tasks with data-driven dependencies is a simple and efficient model for single- and multi-GPU programs. Sequential Task Flow (STF) is such a model that derives task graphs from data dependencies.We propose CUDASTF, a C++ library that implements STF over CUDA APIs, fostering easy creation of scalable and composable algorithms. Users may easily elect to use CUDA Graphs instead of streams, which improves performance of small kernels. Structured kernels are automatically spread over multiple devices and can exercise fine-grained affinity control. Implementationwise, CUDASTF makes a compelling argument for an event-based approach to asynchronous parallel libraries.We obtain up to a 1.8 x improvement over the cuSolverMg library on Cholesky decomposition. On a small weather simulation task we demonstrate near-optimal scalability of our multiGPU kernels; also, on a single GPU, CUDA Graphs improve performance by up to 30%30 \%. Finally, we were able to author the first implementation of the CKKS fully homomorphic encryption scheme over multiple devices.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines