DeepCuts: a deep learning optimization framework for versatile GPU workloads
Wookeun Jung, Thanh Tuan Dao, Jaejin Lee
Abstract
Widely used Deep Learning (DL) frameworks, such as TensorFlow, PyTorch, and MXNet, heavily rely on the NVIDIA cuDNN for performance. However, using cuDNN does not always give the best performance. One reason is that it is hard to handle every case of versatile DNN models and GPU architectures with a library that has a fixed implementation. Another reason is that cuDNN lacks kernel fusion functionality that gives a lot of chances to improve performance. In this paper, we propose a DL optimization framework for versatile GPU workloads, called DeepCuts. It considers both kernel implementation parameters and GPU architectures. It analyzes the DL workload, groups multiple DL operations into a single GPU kernel, and generates optimized GPU kernels considering kernel implementation parameters and GPU architecture parameters. The evaluation result with various DL workloads for inference and training indicates that DeepCuts outperforms cuDNN/cuBLAS-based implementations and the state-of-the-art DL optimization frameworks, such as TVM, TensorFlow XLA, and TensorRT.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 05fa650a-55fc-4edb-80fd-02b752096dcbCited by top-tier papers6
- Welder: Scheduling Deep Learning Memory Access via Tile-graphYining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma et al.OSDI 2023 · 64 citations
- TLP: A Deep Learning-Based Cost Model for Tensor Program TuningYi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu et al.ASPLOS 2023 · 42 citations
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang et al.ASPLOS 2024 · 14 citations
- PolyJuice: Detecting Mis-compilation Bugs in Tensor Compilers with Equality Saturation Based RewritingChijin Zhou, Bingzhou Qian, Gwihwan Go, Quan Zhang et al.OOPSLA 2024 · 7 citations
- REASONING COMPILER: LLM-Guided Optimizations for Efficient Model ServingAnnabelle Sujun Tang, Christopher Priebe, Rohan Mahapatra, Lianhui Qin et al.NeurIPS 2025 · 7 citations
Related papers
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 102 citations
- AutoGraph: Optimizing DNN Computation Graph for Parallel GPU Kernel ExecutionYuxuan Zhao, Qi Sun, Zhuolun He, Yang Bai et al.AAAI 2023 · 10 citations
- DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning WorkloadsQidong Zhao, Hao Wu, Yueming Hao, Zilingfeng Ye et al.ASPLOS 2025 · 3 citations
- Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloadsEvangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman et al.SC 2021 · 2 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
