Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPU
Binqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco Caccamo
Abstract
Pipelining on Edge Tensor Processing Units (TPUs) optimizes the deep neural network (DNN) inference by breaking it down into multiple stages processed concurrently on multiple accelerators. Such DNN inference tasks can be modeled as sporadic non-preemptive gangs with execution times that vary with their parallelism levels. This paper proposes a strict partitioning strategy for deploying DNN inferences in real-time systems. The strategy determines tasks' parallelism levels and assigns tasks to disjoint processor partitions. Configuring the tasks in the same partition with a uniform parallelism level avoids scheduling anomalies and enables schedulability verification using well-understood uniprocessor analyses. Evaluation using real-world Edge TPU benchmarks demonstrated that the proposed method achieves a higher schedulability ratio than state-of-the-art gang scheduling techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5474a438-77e8-4e65-9b27-c9216dac9511Builds on6
- Generating Utilization Vectors for the Systematic Evaluation of Schedulability TestsDavid Griffin, Iain Bate, Robert I. DavisRTSS 2020 · 71 citations
- Design and Timing Guarantee for Non-Preemptive Gang SchedulingSeongtae Lee, Nan Guan, Jinkyu LeeRTSS 2022 · 13 citations
- A Utilization-based Test for Non-preemptive Gang Tasks on MultiprocessorsZheng Dong, Cong LiuRTSS 2022 · 12 citations
- A Universal Method for Task Allocation on FP-FPS Multiprocessor Systems with Spin LocksShuai Zhao, Nan Chen, Yinjie Fang, Zhao Li et al.DAC 2023 · 8 citations
- SPET: Transparent SRAM Allocation and Model Partitioning for Real-time DNN Tasks on Edge TPUChanghun Han, Hoon Sung Chwa, Kilho Lee, Sangeun OhDAC 2023 · 6 citations
Related papers
- Real-Time Multitasking of Deep Neural Networks With Nvidia TensorrtFederico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. ButtazzoRTSS 2025 · 1 citation
- RTInfer: Real-Time Inference of Multiple DNNs on Edge GPUsRenjie Li, Tong Sun, Yi Gao, Wei DongICML 2026
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin et al.RTSS 2021 · 68 citations
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 9 citations
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou et al.INFOCOM 2022 · 28 citations
