Lune

INFOCOM2026顶会

MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile Devices

Zekai Li, Xiaoyi Fan, Xiping Hu, Yifei Zhu

2026年份
1被引次数

摘要

Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖