Real-Time Multitasking of Deep Neural Networks With Nvidia Tensorrt
Federico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. Buttazzo
摘要
Graphics processing units (GPUs) are often employed to accelerate the inference of deep neural networks (DNNs) in cyber-physical systems to implement advanced perception and control functionalities. Frameworks for GPU-accelerated DNN inference typically aim at maximizing the processing throughput rather than focusing on providing a predictable timing behavior, which is crucial for time-sensitive cyber-physical systems. This work proposes a framework for GPU-accelerated inference of DNNs on GPU-based embedded platforms in multitasking scenarios, which provides enhanced timing predictability using a design-time optimization procedure of the DNN workload and a specialized method to schedule the GPU acceleration requests of the DNNs at runtime based on fixed-priority limitedpreemptive scheduling. Fine-grained control of the inference is achieved by splitting the DNNs into smaller chunks, which are then scheduled using a specialized real-time scheduling mechanism. Experimental results on commercial embedded platforms report significant improvements in terms of schedulability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 被引用 153 次
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin 等RTSS 2021 · 被引用 68 次
- On Removing Algorithmic Priority Inversion from Mission-critical Machine Inference PipelinesShengzhong Liu, Shuochao Yao, Xinzhe Fu, Rohan Tabish 等RTSS 2020 · 被引用 53 次
相关 Paper
- SPET: Transparent SRAM Allocation and Model Partitioning for Real-time DNN Tasks on Edge TPUChanghun Han, Hoon Sung Chwa, Kilho Lee, Sangeun OhDAC 2023 · 被引用 6 次
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 被引用 8 次
- RT-MDM: Real-Time Scheduling Framework for Multi-DNN on MCU Using External MemorySukmin Kang, Seongtae Lee, Hyunwoo Koo, Hoon Sung Chwa 等DAC 2024 · 被引用 1 次
- Zygarde: Time-Sensitive On-Device Deep Inference and Adaptation on Intermittently-Powered SystemsBashima Islam, Shahriar NirjonUbiComp 2020 · 被引用 68 次
- Balancing Energy Efficiency and Real-Time Performance in GPU SchedulingYidi Wang, Mohsen Karimi, Yecheng Xiang, Hyoseung KimRTSS 2021 · 被引用 29 次
