MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile Devices
Zekai Li, Xiaoyi Fan, Xiping Hu, Yifei Zhu
Abstract
Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLMKaijie Xiao, Yi Gao, Wei DongUbiComp 2026
- HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live StreamingJiahui Chen, Bo Peng, Lianchen Jia, Zeyu Zhang et al.ICLR 2026 · 1 citation
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu et al.AAAI 2026 · 1 citation
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 69 citations
- MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement LearningYueqian Wang, Songxiang Liu, Disong Wang, Nuo Xu et al.ICLR 2026 · 21 citations
