Bridging the GPU Utilization Gap: Predictive Multi-Dimensional Resource Scheduling for AI Workloads
Yilei Lu, Dongbiao He, Teng Ma, Zhe Liu, Letian Ruan, Jinlei Jiang, Yongwei Wu
摘要
Modern AI data centers face a critical paradox: while machine learning workloads dominate infrastructure demands, actual GPU utilization remains consistently low. Existing schedulers fail to coordinate heterogeneous resources effectively, lack predictive capabilities for dynamic workloads, and cannot balance isolation requirements with sharing optimization in multi-tenant clusters. This paper presents Wind, a novel resource scheduler that bridges the GPU utilization gap through predictive scheduling and geometric resource coordination. Wind introduces three key innovations: (1) a resource prediction framework that leverages historical execution patterns to forecast task requirements and completion times with high accuracy;(2) a unified scheduling architecture supporting isolation, sharing, preemption, and prioritization policies that eliminate resource fragmentation while maintaining performance guarantees; and (3) a Hilbert curve-based multi-dimensional scheduling algorithm that maps CPU-memory-GPU resource space to preserve spatial locality while achieving linear computational complexity.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementJiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang 等ASPLOS 2026 · 被引用 1 次
- Themis: Fair and Efficient GPU Cluster SchedulingKshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman 等NSDI 2020 · 被引用 22 次
- PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU ClustersRutwik Jain, Brandon Tran, Keting Chen, Matthew D. Sinclair 等SC 2024 · 被引用 10 次
- Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUsFahao Chen, Peng Li, Celimuge Wu, Song GuoHPDC 2022 · 被引用 10 次
