LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
Woosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin, Hoon Sung Chwa
摘要
Deep neural networks (DNNs) have shown remarkable success in various machine-learning (ML) tasks useful for many safety-critical, real-time embedded systems. The foremost design goal for enabling DNN execution on real-time embedded systems is to provide worst-case timing guarantees with limited computing resources. Yet, the state-of-the-art ML frameworks hardly leverage heterogeneous computing resources (i.e., CPU, GPU) to improve the schedulability of real-time DNN tasks due to several factors, which include a coarse-grained resource allocation model (one-resource-per-task), the asymmetric nature of DNN execution on CPU and GPU, and lack of schedulability-aware CPU/GPU allocation scheme. This paper presents, to the best of our knowledge, the first study of addressing the above three major barriers and examining their cooperative effect on schedulability improvement. In this paper, we propose LaLaRAND, a real-time layer-level DNN scheduling framework, that enables flexible CPU/GPU scheduling of individual DNN layers by tightly coupling CPU-friendly quantization with fine-grained CPU/GPU allocation schemes (one-resource-per-layer) while mitigating accuracy loss without compromising timing guarantees. We have implemented and evaluated LaLaRAND on top of the state-of-the-art ML framework to demonstrate its effectiveness in making more DNN task sets schedulable by 56% and 80% over an existing approach and a baseline (vanilla PyTorch), respectively, with only up to -0.4% of performance (inference accuracy) difference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous DrivingShuyao Shi, Neiwen Ling, Zhehao Jiang, Xuan Huang 等MobiCom 2024 · 被引用 25 次
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn 等RTSS 2022 · 被引用 21 次
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 被引用 10 次
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio 等RTSS 2023 · 被引用 9 次
- RED: A Systematic Real-Time Scheduling Approach for Robotic Environmental DynamicsZexin Li, Tao Ren, Xiaoxi He, Cong LiuRTSS 2023 · 被引用 8 次
它引用的顶会 Paper4
- SpiroSonic: monitoring human lung function via acoustic sensing on commodity smartphonesXingzhe Song, Boyuan Yang, Ge Yang, Ruirong Chen 等MobiCom 2020 · 被引用 97 次
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 被引用 71 次
- On Removing Algorithmic Priority Inversion from Mission-critical Machine Inference PipelinesShengzhong Liu, Shuochao Yao, Xinzhe Fu, Rohan Tabish 等RTSS 2020 · 被引用 53 次
- R-TOD: Real-Time Object Detector with Minimized End-to-End Delay for Autonomous DrivingWonseok Jang, Hansaem Jeong, Kyungtae Kang, Nikil D. Dutt 等RTSS 2020 · 被引用 37 次
相关 Paper
- SPET: Transparent SRAM Allocation and Model Partitioning for Real-time DNN Tasks on Edge TPUChanghun Han, Hoon Sung Chwa, Kilho Lee, Sangeun OhDAC 2023 · 被引用 6 次
- Real-Time Multitasking of Deep Neural Networks With Nvidia TensorrtFederico Aromolo, Andrea Stevanato, Alessandro Biondi, Giorgio C. ButtazzoRTSS 2025 · 被引用 1 次
- Zygarde: Time-Sensitive On-Device Deep Inference and Adaptation on Intermittently-Powered SystemsBashima Islam, Shahriar NirjonUbiComp 2020 · 被引用 68 次
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 被引用 4 次
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 被引用 8 次
