CASE: a compiler-assisted SchEduling framework for multi-GPU systems
Chao Chen, Chris Porter, Santosh Pande
Abstract
Modern computing platforms tend to deploy multiple GPUs on a single node to boost performance. GPUs have large computing capacities and are an expensive resource. Increasing their utilization without causing performance degradation of individual workloads is an important and challenging problem. Although services such as NVIDIA's MPS allow multiple cooperative kernels to simultaneously run on a single device, they do not solve the co-execution problem for uncooperative, independent kernels on such a multi-GPU system. To tackle this problem, we propose CASE --- a fully automated compiler-assisted scheduling framework. During the compilation of an application, CASE constructs GPU tasks from CUDA programs and instruments the code with a probe before each one. At runtime, each probe conveys information about its task's resource requirements such as memory and the number of streaming multiprocessor (SMs) needed to a user-level scheduler. The scheduler then places each task onto a suitable device by employing a policy appropriate to the system. In our prototype, a throughput-oriented scheduling policy is implemented to evaluate our resource-aware scheduling framework. The Rodinia benchmark suite and the Darknet neural network framework were used in our evaluation. The results show that, as compared to existing state-of-the-art methods, CASE improves throughput by up to 2.5X for Rodinia, and up to 2.7X for Darknet on modern NVIDIA GPU platforms, mainly due to the fact that it improves the average system utilization by up to 3.36X and the job turnaround time by up to 4.9X. Meanwhile, it limits individual kernel performance degradation within 2.5%. CASE achieved peak system utilization of 78% for Rodinia and 80% for Darknet on a 4XV100 system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbaa57cb-2659-4591-9038-4a1656f5408dCited by top-tier papers2
- Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning WorkloadsWei Zhao, Anand Jayarajan, Gennady PekhimenkoASPLOS 2025 · 4 citations
- Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal OrchestrationZejia Lin, Hongxin Xu, Guanyi Chen, Zhiguang Chen et al.ASPLOS 2026 · 2 citations
Related papers
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 4 citations
- Deadline-Aware Offloading for High-Throughput AcceleratorsTsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. RogersHPCA 2021 · 16 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
- Real-Time Multitasking of Deep Neural Networks With Nvidia TensorrtFederico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. ButtazzoRTSS 2025 · 1 citation
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
