Real-Time Multitasking of Deep Neural Networks With Nvidia Tensorrt
Federico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. Buttazzo
Abstract
Graphics processing units (GPUs) are often employed to accelerate the inference of deep neural networks (DNNs) in cyber-physical systems to implement advanced perception and control functionalities. Frameworks for GPU-accelerated DNN inference typically aim at maximizing the processing throughput rather than focusing on providing a predictable timing behavior, which is crucial for time-sensitive cyber-physical systems. This work proposes a framework for GPU-accelerated inference of DNNs on GPU-based embedded platforms in multitasking scenarios, which provides enhanced timing predictability using a design-time optimization procedure of the DNN workload and a specialized method to schedule the GPU acceleration requests of the DNNs at runtime based on fixed-priority limitedpreemptive scheduling. Fine-grained control of the inference is achieved by splitting the DNNs into smaller chunks, which are then scheduled using a specialized real-time scheduling mechanism. Experimental results on commercial embedded platforms report significant improvements in terms of schedulability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8e4e3d0-a275-44bb-9945-0eccefd156fdBuilds on4
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin et al.RTSS 2021 · 68 citations
- On Removing Algorithmic Priority Inversion from Mission-critical Machine Inference PipelinesShengzhong Liu, Shuochao Yao, Xinzhe Fu, Rohan Tabish et al.RTSS 2020 · 53 citations
Related papers
- SPET: Transparent SRAM Allocation and Model Partitioning for Real-time DNN Tasks on Edge TPUChanghun Han, Hoon Sung Chwa, Kilho Lee, Sangeun OhDAC 2023 · 6 citations
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 8 citations
- RT-MDM: Real-Time Scheduling Framework for Multi-DNN on MCU Using External MemorySukmin Kang, Seongtae Lee, Hyunwoo Koo, Hoon Sung Chwa et al.DAC 2024 · 1 citation
- Zygarde: Time-Sensitive On-Device Deep Inference and Adaptation on Intermittently-Powered SystemsBashima Islam, Shahriar NirjonUbiComp 2020 · 68 citations
- Balancing Energy Efficiency and Real-Time Performance in GPU SchedulingYidi Wang, Mohsen Karimi, Yecheng Xiang, Hyoseung KimRTSS 2021 · 29 citations
