Server-Driven Video Streaming for Deep Learning Inference
Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, Junchen Jiang
摘要
Video streaming is crucial for AI applications that gather videos from sources to servers for inference by deep neural nets (DNNs). Unlike traditional video streaming that optimizes visual quality, this new type of video streaming permits aggressive compression/pruning of pixels not relevant to achieving high DNN inference accuracy. However, much of this potential is left unrealized, because current video streaming protocols are driven by the video source (camera) where the compute is rather limited. We advocate that the video streaming protocol should be driven by real-time feedback from the server-side DNN. Our insight is two-fold: (1) server-side DNN has more context about the pixels that maximize its inference accuracy; and (2) the DNN's output contains rich information useful to guide video streaming. We present DDS (DNN-Driven Streaming), a concrete design of this approach. DDS continuously sends a low-quality video stream to the server; the server runs the DNN to determine where to re-send with higher quality to increase the inference accuracy. We find that compared to several recent baselines on multiple video genres and vision tasks, DDS maintains higher accuracy while reducing bandwidth usage by upto 59% or improves accuracy by upto 9% with no additional bandwidth usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan 等NSDI 2023 · 被引用 94 次
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 被引用 69 次
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 被引用 64 次
- Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine LearningChengfei Lv, Chaoyue Niu, Renjie Gu, Xiaotang Jiang 等OSDI 2022 · 被引用 52 次
它引用的顶会 Paper1
相关 Paper
- AdaMask: Enabling Machine-Centric Video Streaming with Adaptive Frame Masking for DNN Inference OffloadingShengzhong Liu, Tianshi Wang, Jinyang Li, Dachun Sun 等ACM MM 2022 · 被引用 45 次
- VidIQ: Inference-Aware Neural Codecs for Quality-Enhanced, Real-Time Video AnalyticsAndong Zhu, Sheng Zhang, Xiaohang Shi, Hesheng Sun 等ACM MM 2025
- AdaStreamer: Machine-Centric High-Accuracy Multi-Video Analytics with Adaptive Neural CodecsAndong Zhu, Sheng Zhang, Ke Cheng, Xiaohang Shi 等INFOCOM 2024 · 被引用 8 次
- Batch Adaptative Streaming for Video AnalyticsLei Zhang, Yuqing Zhang, Ximing Wu, Fangxin Wang 等INFOCOM 2022 · 被引用 24 次
- An Intelligent Video Processing Architecture for Edge-cloud Video StreamingChengsi Gao, Ying Wang, Weiwei Chen, Lei ZhangDAC 2021 · 被引用 6 次
