OmniBoost: Boosting Throughput of Heterogeneous Embedded Devices under Multi-DNN Workload
Andreas Karatzas, Iraklis Anagnostopoulos
Abstract
Modern Deep Neural Networks (DNNs) exhibit profound efficiency and accuracy properties. This has introduced application workloads that comprise of multiple DNN applications, raising new challenges regarding workload distribution. Equipped with a diverse set of accelerators, newer embedded system present architectural heterogeneity, which current run-time controllers are unable to fully utilize. To enable high throughput in multi-DNN workloads, such a controller is ought to explore hundreds of thousands of possible solutions to exploit the underlying heterogeneity. In this paper, we propose OmniBoost, a lightweight and extensible multi-DNN manager for heterogeneous embedded devices. We leverage stochastic space exploration and we combine it with a highly accurate performance estimator to observe a ×4.6 average throughput boost compared to other state-of-the-art methods. The evaluation was performed on the HiKey970 development board. Our code is publicly available at https://github.com/AndreasKaratzas/omniboost-v1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59a5424c-774f-43fb-8dd1-c9dc04b5293fCited by top-tier papers2
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 18 citations
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han et al.DAC 2025
Builds on2
Related papers
- NeuOS: A Latency-Predictable Multi-Dimensional Optimization Framework for DNN-driven Autonomous SystemsSoroush Bateni, Cong LiuUSENIX ATC 2020 · 49 citations
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
- Centimani: Enabling Fast AI Accelerator Selection for DNN Training with a Novel Performance PredictorZhen Xie, Murali Emani, Xiaodong Yu, Dingwen Tao et al.USENIX ATC 2024 · 4 citations
- Decentralized Application-Level Adaptive Scheduling for Multi-Instance DNNs on Open Mobile DevicesHsin-Hsuan Sung, Jou-An Chen, Wei Niu, Jiexiong Guan et al.USENIX ATC 2023 · 9 citations
- DREAM: A Dynamic Scheduler for Dynamic Real-time Multi-model ML WorkloadsSeah Kim, Hyoukjun Kwon, Jinook Song, Jihyuck Jo et al.ASPLOS 2023 · 20 citations
