RTInfer: Real-Time Inference of Multiple DNNs on Edge GPUs
Renjie Li, Tong Sun, Yi Gao, Wei Dong
Abstract
While edge GPUs are increasingly used for latency-critical DNN tasks, limited resources often fail to meet strict real-time (RT) requirements under concurrent workloads. Existing preemption and early-exit mechanisms often underutilize GPU resources through single-task queuing and sacrifice excessive accuracy during task bursts. To address this, we propose RTInfer, a novel system that enables concurrent RT task execution while balancing throughput and accuracy. RTInfer integrates an accuracy-calibrated lightweight variant co-optimization to generate efficient models, a memory-layout-aware scheduler to mitigate fragmentation during preemption, and an on-demand loading strategy to minimize host-to-GPU latency. Extensive evaluations demonstrate that RTInfer outperforms state-of-the-art methods by reducing average deadline miss rate (DMR) from 32.8% to 0% and improving accuracy by up to 56.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4f25212-57bd-4d70-8180-500155321361Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Width & Depth Pruning for Vision TransformersFang Yu, Kun Huang, Meng Wang, Yuan Cheng et al.AAAI 2022 · 159 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin et al.ASPLOS 2020 · 143 citations
Related papers
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 8 citations
- SPET: Transparent SRAM Allocation and Model Partitioning for Real-time DNN Tasks on Edge TPUChanghun Han, Hoon Sung Chwa, Kilho Lee, Sangeun OhDAC 2023 · 6 citations
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin et al.RTSS 2021 · 68 citations
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou et al.INFOCOM 2022 · 28 citations
- ElasticRoom: Multi-Tenant DNN Inference Engine via Co-design with Resource-constrained Compilation and Strong Priority SchedulingLixian Ma, Haoruo Chen, En Shao, Leping Wang et al.HPDC 2024 · 4 citations
