Real-Time Video Inference on Edge Devices via Adaptive Model Streaming
Mehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad Alizadeh
摘要
Real-time video inference on edge devices like mobile phones and drones is challenging due to the high computation cost of Deep Neural Networks. We present Adaptive Model Streaming (AMS), a new approach to improving performance of efficient lightweight models for video inference on edge devices. AMS uses a remote server to continually train and adapt a small model running on the edge device, boosting its performance on the live video using online knowledge distillation from a large, state-of-the-art model. We discuss the challenges of over-the-network model adaptation for video inference, and present several techniques to reduce communication cost of this approach: avoiding excessive overfitting, updating a small fraction of important model parameters, and adaptive sampling of training frames at edge devices. On the task of video semantic segmentation, our experimental results show 0.4-17.8 percent mean Intersection-over-Union improvement compared to a pretrained model across several video datasets. Our prototype can perform video segmentation at 30 frames-per-second with 40 milliseconds camera-to-label latency on a Samsung Galaxy S10+ mobile phone, using less than 300 Kbps uplink and downlink bandwidth on the device. No Customization 0 Up / 0 Down (Kbps) One-Time Customized 50 Up / 80 Down (Kbps) Adaptive Model Streaming (Ours) 169 Up / 206 Down (Kbps) Remote Inference + Optical Flow Tracking 1880 Up / 30 Down (Kbps) Road Building Vegetation Sky Person Car Just-In-Time 2520 Up / 3207 Down (Kbps)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Shoggoth: Towards Efficient Edge-Cloud Collaborative Real-Time Video Inference via Adaptive Online LearningLiang Wang, Kai Lu, Nan Zhang, Xiaoyang Qu 等DAC 2023 · 被引用 25 次
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim 等ISCA 2024 · 被引用 13 次
- Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeLiang Wang, Xiaoyang Qu, Jianzong Wang, Guokuan Li 等INFOCOM 2024 · 被引用 11 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
- Multi-Edge Reinforced Collaborative Data Acquisition for Continuous Video Analytics by Prioritizing Quality over QuantityLei Zhang, Guanyu Gao, Haiyan Yin, Huaizheng ZhangAAAI 2025 · 被引用 2 次
它引用的顶会 Paper3
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Online Model Distillation for Efficient Video InferenceRavi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan 等ICCV 2019 · 被引用 131 次
- SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationXianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi 等CVPR 2020
相关 Paper
- DeDelayed: Deleting Remote Inference Delay via On-Device CorrectionDan Jacobellis, Mateen Ulhaq, Fabien Racapé, Hyomin Choi 等CVPR 2026 · 被引用 2 次
- AdaMask: Enabling Machine-Centric Video Streaming with Adaptive Frame Masking for DNN Inference OffloadingShengzhong Liu, Tianshi Wang, Jinyang Li, Dachun Sun 等ACM MM 2022 · 被引用 45 次
- MobileFaceSwap: A Lightweight Framework for Video Face SwappingZhiliang Xu, Zhibin Hong, Changxing Ding, Zhen Zhu 等AAAI 2022 · 被引用 78 次
- Real-Time, Accurate, and Consistent Video Semantic Segmentation via Unsupervised Adaptation and Cross-Unit Deployment on Mobile DeviceHyojin Park, Alan Yessenbayev, Tushar Singhal, Navin Kumar Adhikari 等CVPR 2022 · 被引用 8 次
- Edge-Assisted On-Device Model Update for Video Analytics in Adverse EnvironmentsYuxin Kong, Peng Yang, Yan ChengACM MM 2023 · 被引用 40 次
