Deadline-Aware Offloading for High-Throughput Accelerators
Tsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. Rogers
摘要
Contemporary GPUs are widely used for throughput-oriented data-parallel workloads and increasingly are being considered for latency-sensitive applications in datacenters. Examples include recurrent neural network (RNN) inference, network packet processing, and intelligent personal assistants. These data parallel applications have both high throughput demands and real-time deadlines (40µs-7ms). Moreover, the kernels in these applications have relatively few threads that do not fully utilize the device unless a large batch size is used. However, batching forces jobs to wait, which increases their latency, especially when realistic job arrival times are considered.
Previously, programmers have managed the tradeoffs associated with concurrent, latency-sensitive jobs by using a combination of GPU streams and advanced scheduling algorithms running on the CPU host. Although GPU streams allow the accelerator to execute multiple jobs concurrently, prior state-of-the-art solutions use the relatively distant CPU host to prioritize the latency-sensitive GPU tasks. Thus, these approaches are forced to operate at a coarse granularity and cannot quickly adapt to rapidly changing program behavior.
We observe that fine-grain, device-integrated kernel schedulers efficiently meet the deadlines of concurrent, latencysensitive GPU jobs. To overcome the limitations of softwareonly, CPU-side approaches, we extend the GPU queue scheduler to manage real-time deadlines. We propose a novel laxity-aware scheduler (LAX) that uses information collected within the GPU to dynamically vary job priority based on how much laxity jobs have before their deadline. Compared to contemporary GPUs, 3 state-of-the-art CPU-side schedulers and 6 other advanced GPU-side schedulers, LAX meets the deadlines of 1.7X -5.0X more jobs and provides better energy-efficiency, throughput, and 99-percentile tail latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 被引用 153 次
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen 等OSDI 2025 · 被引用 9 次
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 被引用 3 次
它引用的顶会 Paper2
相关 Paper
- ElasticRoom: Multi-Tenant DNN Inference Engine via Co-design with Resource-constrained Compilation and Strong Priority SchedulingLixian Ma, Haoruo Chen, En Shao, Leping Wang 等HPDC 2024 · 被引用 4 次
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 被引用 4 次
- CASE: a compiler-assisted SchEduling framework for multi-GPU systemsChao Chen, Chris Porter, Santosh PandePPoPP 2022 · 被引用 16 次
- Tacker: Tensor-CUDA Core Kernel Fusion for Improving the GPU Utilization while Ensuring QoSHan Zhao, Weihao Cui, Quan Chen, Youtao Zhang 等HPCA 2022 · 被引用 42 次
- RELIEF: Relieving Memory Pressure In SoCs Via Data Movement-Aware Accelerator SchedulingSudhanshu Gupta, Sandhya DwarkadasHPCA 2024 · 被引用 4 次
