Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloading
Wuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia, Yunxin Liu, Marco Gruteser, Dipankar Raychaudhuri, Yanyong Zhang
摘要
As mobile devices continuously generate streams of images and videos, a new class of mobile deep vision applications are rapidly emerging, which usually involve running deep neural networks on these multimedia data in real-time. To support such applications, having mobile devices offload the computation, especially the neural network inference, to edge clouds has proved effective. Existing solutions often assume there exists a dedicated and powerful server, to which the entire inference can be offloaded. In reality, however, we may not be able to find such a server but need to make do with less powerful ones. To address these more practical situations, we propose to partition the video frame and offload the partial inference tasks to multiple servers for parallel processing. This paper presents the design of Elf, a framework to accelerate the mobile deep vision applications with any server provisioning through the parallel offloading. Elf employs a recurrent region proposal prediction algorithm, a region proposal centric frame partitioning, and a resource-aware multi-offloading scheme. We implement and evaluate Elf upon Linux and Android platforms using four commercial mobile devices and three deep vision applications with ten state-of-the-art models. The comprehensive experiments show that Elf can speed up the applications by 4.85× with saving bandwidth usage by 52.6%, while with <1% application accuracy sacrifice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu 等MobiCom 2021 · 被引用 103 次
- Real-time neural network inference on extremely weak devices: agile offloading with explainable AIKai Huang, Wei GaoMobiCom 2022 · 被引用 57 次
- MergeSFL: Split Federated Learning with Feature Merging and Batch Size RegulationYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等ICDE 2024 · 被引用 41 次
- Edge-Assisted On-Device Model Update for Video Analytics in Adverse EnvironmentsYuxin Kong, Peng Yang, Yan ChengACM MM 2023 · 被引用 40 次
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang 等MobiCom 2024 · 被引用 40 次
它引用的顶会 Paper6
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- A First Look at Commercial 5G Performance on SmartphonesArvind Narayanan, Eman Ramadan, Jason Carpenter, Qingxu Liu 等WWW 2020 · 被引用 268 次
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang 等SIGCOMM 2020 · 被引用 264 次
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery 等SIGCOMM 2020 · 被引用 238 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
相关 Paper
- ResMap: Exploiting Sparse Residual Feature Map for Accelerating Cross-Edge Video AnalyticsNing Chen, Shuai Zhang, Sheng Zhang, Yuting Yan 等INFOCOM 2023 · 被引用 11 次
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 被引用 213 次
- Resource-aware Deployment of Dynamic DNNs over Multi-tiered Interconnected SystemsChetna Singhal, Yashuo Wu, Francesco Malandrino, Marco Levorato 等INFOCOM 2024 · 被引用 15 次
- Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online LearningLetian Zhang, Lixing Chen, Jie XuWWW 2021 · 被引用 75 次
- EdgeDuet: Tiling Small Object Detection for Edge Assisted Autonomous Mobile VisionXu Wang, Zheng Yang, Jiahang Wu, Yi Zhao 等INFOCOM 2021 · 被引用 54 次
