Server-Driven Video Streaming for Deep Learning Inference
Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, Junchen Jiang
Abstract
Video streaming is crucial for AI applications that gather videos from sources to servers for inference by deep neural nets (DNNs). Unlike traditional video streaming that optimizes visual quality, this new type of video streaming permits aggressive compression/pruning of pixels not relevant to achieving high DNN inference accuracy. However, much of this potential is left unrealized, because current video streaming protocols are driven by the video source (camera) where the compute is rather limited. We advocate that the video streaming protocol should be driven by real-time feedback from the server-side DNN. Our insight is two-fold: (1) server-side DNN has more context about the pixels that maximize its inference accuracy; and (2) the DNN's output contains rich information useful to guide video streaming. We present DDS (DNN-Driven Streaming), a concrete design of this approach. DDS continuously sends a low-quality video stream to the server; the server runs the DNN to determine where to re-send with higher quality to increase the inference accuracy. We find that compared to several recent baselines on multiple video genres and vision tasks, DDS maintains higher accuracy while reducing bandwidth usage by upto 59% or improves accuracy by upto 9% with no additional bandwidth usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan et al.NSDI 2023 · 94 citations
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 69 citations
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 64 citations
- Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine LearningChengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang et al.OSDI 2022 · 52 citations
Builds on1
Related papers
- AdaMask: Enabling Machine-Centric Video Streaming with Adaptive Frame Masking for DNN Inference OffloadingShengzhong Liu, Tianshi Wang, Jinyang Li, Dachun Sun et al.ACM MM 2022 · 45 citations
- VidIQ: Inference-Aware Neural Codecs for Quality-Enhanced, Real-Time Video AnalyticsAndong Zhu, Sheng Zhang, Xiaohang Shi, Hesheng Sun et al.ACM MM 2025
- AdaStreamer: Machine-Centric High-Accuracy Multi-Video Analytics with Adaptive Neural CodecsAndong Zhu, Sheng Zhang, Ke Cheng, Xiaohang Shi et al.INFOCOM 2024 · 8 citations
- Batch Adaptative Streaming for Video AnalyticsLei Zhang, Yuqing Zhang, Ximing Wu, Fangxin Wang et al.INFOCOM 2022 · 24 citations
- An Intelligent Video Processing Architecture for Edge-cloud Video StreamingChengsi Gao, Ying Wang, Weiwei Chen, Lei ZhangDAC 2021 · 6 citations
