NeuOS: A Latency-Predictable Multi-Dimensional Optimization Framework for DNN-driven Autonomous Systems
Soroush Bateni, Cong Liu
Abstract
Deep neural networks (DNNs) used in computer vision have become widespread techniques commonly used in autonomous embedded systems for applications such as image/object recognition and tracking. The stringent space, weight, and power constraints seen in such systems impose a major impediment for practical and safe implementation of DNNs, because they have to be latency predictable while ensuring minimum energy consumption and maximum accuracy. Unfortunately, exploring this optimization space is very challenging because (1) smart coordination has to be performed among system-and application-level solutions, (2) layer characteristics should be taken into account, and more importantly, (3) when multiple DNNs exist, a consensus on system configurations should be calculated, which is a problem that is an order of magnitude harder than any previously considered scenario. In this paper, we present NeuOS, a comprehensive latency predictable system solution for running multi-DNN workloads in autonomous systems. NeuOS can guarantee latency predictability, while managing energy optimization and dynamic accuracy adjustment based on specific system constraints via smart coordinated systemand application-level decision-making among multiple DNN instances. We implement and extensively evaluate NeuOS on two state-of-the-art autonomous system platforms for a set of popular DNN models. Experiments show that NeuOS rarely misses deadlines, and can improve energy and accuracy considerably compared to state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c59d1f34-a125-410b-8d1a-3e19f363cf37Cited by top-tier papers11
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLZiqi Zhang, Chen Gong, Yifeng Cai, Yuanyuan Yuan et al.S&P 2024 · 53 citations
- A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile DevicesChengdong Lin, Kun Wang, Zhenjiang Li, Yu PuMobiCom 2023 · 52 citations
- LegoDNN: block-grained scaling of deep neural networks for mobile visionRui Han, Qinglong Zhang, Chi Harold Liu, Guoren Wang et al.MobiCom 2021 · 51 citations
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li et al.USENIX ATC 2025 · 41 citations
- DeepPerform: An Efficient Approach for Performance Testing of Resource-Constrained Neural NetworksSimin Chen, Mirazul Haque, Cong Liu, Wei YangASE 2022 · 19 citations
Related papers
- ALERT: Accurate Learning for Energy and TimelinessChengcheng Wan, Muhammad Husni Santriaji, Eri Rogers, Henry Hoffmann et al.USENIX ATC 2020 · 15 citations
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang et al.INFOCOM 2025 · 2 citations
- AxoNN: energy-aware execution of neural network inference on multi-accelerator heterogeneous SoCsIsmet Dagli, Alexander Cieslewicz, Jedidiah McClurg, Mehmet E. BelviranliDAC 2022 · 38 citations
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 18 citations
- Zygarde: Time-Sensitive On-Device Deep Inference and Adaptation on Intermittently-Powered SystemsBashima Islam, Shahriar NirjonUbiComp 2020 · 68 citations
