MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile Devices
Zekai Li, Xiaoyi Fan, Xiping Hu, Yifei Zhu
摘要
Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLMKaijie Xiao, Yi Gao, Wei DongUbiComp 2026
- HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live StreamingJiahui Chen, Bo Peng, Lianchen Jia, Zeyu Zhang 等ICLR 2026 · 被引用 1 次
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu 等AAAI 2026 · 被引用 1 次
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 被引用 69 次
- MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement LearningYueqian Wang, Songxiang Liu, Disong Wang, Nuo Xu 等ICLR 2026 · 被引用 21 次
