Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUs
Bojian Zheng, Cody Hao Yu, Jie Wang, Yaoyao Ding, Yizhi Liu, Yida Wang, Gennady Pekhimenko
Abstract
Achieving high performance in machine learning workloads is a crucial yet difficult task. To achieve high runtime performance on hardware platforms such as GPUs, graph-based executions such as CUDA graphs are often used to eliminate CPU runtime overheads by submitting jobs in the granularity of multiple kernels. However, many machine learning workloads, especially dynamic deep neural networks (DNNs) with varying-sized inputs or data-dependent control flows, face challenges when directly using CUDA graphs to achieve optimal performance. We observe that the use of graph-based executions poses three key challenges in terms of efficiency and even practicability: (1) Extra data movements when copying input values to graphs’ placeholders. (2) High GPU memory consumption due to the numerous CUDA graphs created to efficiently support dynamic-shape workloads. (3) Inability to handle data-dependent control flows.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8f5cf2e8-8941-40d8-a08f-0506b6afe276Cited by top-tier papers1
Ask how each one uses itRelated papers
- ED-Batch: Efficient Automatic Batching of Dynamic Neural Networks via Learned Finite State MachinesSiyuan Chen, Pratik Pramod Fegade, Tianqi Chen, Phillip B. Gibbons et al.ICML 2023 · 1 citation
- DyCL: Dynamic Neural Network Compilation Via Program Rewriting and Graph OptimizationSimin Chen, Shiyi Wei, Cong Liu, Wei YangISSTA 2023 · 11 citations
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen et al.HPCA 2025 · 2 citations
- MAGPY: Compiling Eager Mode DNN Programs by Monitoring Execution StatesChen Zhang, Rongchao Dong, Haojie Wang, Runxin Zhong et al.USENIX ATC 2024 · 5 citations
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
