SC2024Top-tier venue
CUDASTF: Bridging the Gap Between CUDA and Task Parallelism
Cédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael Garland
Abstract
Organizing computation as asynchronous tasks with data-driven dependencies is a simple and efficient model for single- and multi-GPU programs. Sequential Task Flow (STF) is such a model that derives task graphs from data dependencies.We propose CUDASTF, a C++ library that implements STF over CUDA APIs, fostering easy creation of scalable and composable algorithms. Users may easily elect to use CUDA Graphs instead of streams, which improves performance of small kernels. Structured kernels are automatically spread over multiple devices and can exercise fine-grained affinity control. Implementationwise, CUDASTF makes a compelling argument for an event-based approach to asynchronous parallel libraries.We obtain up to a 1.8 x improvement over the cuSolverMg library on Cholesky decomposition. On a small weather simulation task we demonstrate near-optimal scalability of our multiGPU kernels; also, on a single GPU, CUDA Graphs improve performance by up to . Finally, we were able to author the first implementation of the CKKS fully homomorphic encryption scheme over multiple devices.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 8 citations
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- Streaming Task Graph Scheduling for Dataflow ArchitecturesTiziano De Matteis, Lukas Gianinazzi, Johannes de Fine Licht, Torsten HoeflerHPDC 2023 · 3 citations
- Cheddar: A Swift Fully Homomorphic Encryption Library Designed for GPU ArchitecturesWonseok Choi, Jongmin Kim, Jung Ho AhnASPLOS 2026 · 6 citations
- Choosing the Best Parallelization and Implementation Styles for Graph Analytics Codes: Lessons Learned from 1106 ProgramsYiqian Liu, Noushin Azami, Avery Vanausdal, Martin BurtscherSC 2023 · 4 citations
