UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLM
Kaijie Xiao, Yi Gao, Wei Dong
Abstract
Video analytics is ubiquitous in modern society, and the emergence of Multimodal Large Language Models (MLLMs) has made its application even more extensive. A common way to support MLLM-based video analytics is to continuously stream video to the cloud and then extracts visual features and samples frames for analysis. However, this cloud-centric architecture can introduce high transmission overhead. Simultaneously, the fixed-rate or uniform sampling approaches employed in these systems will miss critical visual evidence relevant to the query, resulting in reduced accuracy.; AB@In this paper, we propose UniOVA , an edge-cloud collaborative framework for universal on-demand video analytics. First, query-aware feature retrieval improves accuracy and reduces transmission overhead by retrieving and transmitting only relevant ViT features from the edge. Second, Interest-Aware ViT reduces edge overhead through hierarchical token merging which compresses irrelevant visual data based on user interests. Our evaluation shows that UniOVA achieves comparable or higher accuracy than LongVA while reducing transmission overhead by 37.1%-95.8%. Compared with ChatCam, UniOVA also improves accuracy by 63.2%-193.8%. In addition, on Jetson Xavier NX, UniOVA achieves 7.14-11.43 FPS for feature extraction, and its end-to-end query latency ranges from 3.67-8.18s under different bandwidth settings. A real-world user study further demonstrates its potential for efficient and precise video analytics tailored to individual user needs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- MoVi: Real-Time Large Multimodal Model-Driven Interactive Video Analytics on Mobile DevicesZekai Li, Xiaoyi Fan, Xiping Hu, Yifei ZhuINFOCOM 2026 · 1 citation
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu et al.ICML 2025
- Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesHaocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu et al.AAAI 2026 · 1 citation
- Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language ModelsXuyang Liu, Yiyu Wang, Junpeng Ma, Linfeng ZhangEMNLP 2025 · 3 citations
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video UnderstandingMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangNeurIPS 2025 · 62 citations
