GraCE: Unlocking CUDA Graphs with Compiler Support for ML Workloads
Abhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava Basu
Abstract
Machine learning (ML) workloads launch hundreds of short-running GPU kernels per iteration. With GPU compute throughput growing rapidly, CPU-side launch latency of kernels is emerging as a bottleneck. CUDA Graphs promise to address this by replaying a set of kernels with a single dispatch of the graph, removing per-kernel launch costs. However, CUDA Graphs remain surprisingly difficult to deploy correctly and efficiently.
We present GraCE -a compiler framework to maximize the coverage and benefits of CUDA Graphs for ML workloads. It introduces three novel optimizations: it applies automatic code transformations to make ML applications amenable for CUDA Graphs; it eliminates the parameter copy overheads for kernels executing in CUDA Graphs, and it selectively deploys CUDA Graphs guided by a cost-benefit analysis. For 25 ML workloads from TorchBench, HuggingFace and TIMM, GraCE more than doubles the benefit from deploying CUDA Graph compared to the most popular and widely used ML compiler PyTorch2. GraCE is built atop PyTorch2's compilation framework and requires no programmer intervention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 710527bd-9bcf-4589-8c67-0bec9f8d59fbBuilds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu et al.OSDI 2020 · 551 citations
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementWonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong SimOSDI 2024 · 248 citations
Related papers
- Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUsBojian Zheng, Cody Hao Yu, Jie Wang, Yaoyao Ding et al.MICRO 2023 · 4 citations
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 8 citations
- AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel ExecutionYuxuan Zhao, Qi Sun, Zhuolun He, Yang Bai et al.AAAI 2023 · 10 citations
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 27 citations
- MPK: A Compiler and Runtime for Mega-Kernelizing Tensor ProgramsXinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji et al.OSDI 2026 · 20 citations
