The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSL
Bo Qiao, M. Akif Özkan, Jürgen Teich, Frank Hannig
摘要
CUDA graph is an asynchronous task-graph programming model recently released by Nvidia. It encapsulates application workflows in a graph, with nodes being operations connected by dependencies. The new API brings two benefits: Reduced work launch overhead and whole workflow optimizations. In this paper, we improve the ability of CUDA graph to exploit workflow optimizations, e.g., concurrent kernel executions with complementary resource occupancy. Additionally, we argue that the advantages of DSLs are complementary to CUDA graph, and joining the two techniques can benefit from the best of both worlds. Here, we propose a compiler-based approach that combines CUDA graph with an image processing DSL and a source-to-source compiler called Hipacc. For ten image processing applications benchmarked on two Nvidia GPUs, our approach is able to achieve a geometric mean speedup of 1.30 over Hipacc without CUDA graph, 1.11 over CUDA graph without Hipacc, and 3.96 over another state-of-the-art DSL called Halide.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 被引用 5 次
- Efficient automatic scheduling of imaging and vision pipelines for the GPULuke Anderson, Andrew Adams, Karima Ma, Tzu-Mao Li 等OOPSLA 2021 · 被引用 14 次
- AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel ExecutionYuxuan Zhao, Qi Sun, Zhuolun He, Yang Bai 等AAAI 2023 · 被引用 10 次
- CUDASTF: Bridging the Gap Between CUDA and Task ParallelismCédric Augonnet, Andrei Alexandrescu, Albert Sidelnik, Michael GarlandSC 2024 · 被引用 7 次
