AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile Web
Yakun Huang, Xiuquan Qiao, Schahram Dustdar, Yan Li
摘要
Employing today’s deep neural network (DNN) into the cross-platform web with an offloading way has been a promising means to alleviate the tension between intensive inference and limited computing resources. However, it is still challenging to directly leverage the distributed DNN execution into web apps with the following limitations, including (1) how special computing tasks such as DNN inference can provide fine-grained and efficient offloading in the inefficient JavaScript-based environment? (2) lacking the ability to balance the latency and mobile energy to partition the inference facing various web applications’ requirements. (3) and ignoring that DNN inference is vulnerable to the operating environment and mobile devices’ computing capability, especially dedicated web apps. This paper designs AoDNN, an automatic offloading framework to orchestrate the DNN inference across the mobile web and the edge server, with three main contributions. First, we design the DNN offloading based on providing a snapshot mechanism and use multi-threads to monitor dynamic contexts, partition decision, trigger offloading, etc. Second, we provide a learning-based latency and mobile energy prediction framework for supporting various web browsers and platforms. Third, we establish a multi-objective optimization to solve the optimal partition by balancing the latency and mobile energy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo 等UbiComp 2020 · 被引用 77 次
- DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningYakun Huang, Xiuquan Qiao, Jian Tang, Pei Ren 等INFOCOM 2020 · 被引用 32 次
相关 Paper
- AccuMO: Accuracy-Centric Multitask Offloading in Edge-Assisted Mobile Augmented RealityZ. Jonny Kong, Qiang Xu, Jiayi Meng, Y. Charlie HuMobiCom 2023 · 被引用 21 次
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato 等INFOCOM 2024 · 被引用 15 次
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 被引用 213 次
- Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online LearningLetian Zhang, Lixing Chen, Jie XuWWW 2021 · 被引用 75 次
- Deep Learning on Mobile Devices Through Neural Processing Units and Edge ComputingTianxiang Tan, Guohong CaoINFOCOM 2022 · 被引用 35 次
