Artic: AI-oriented Real-time Communication for MLLM Video Assistant
Jiangkai Wu, Zhiyuan Ren, Junquan Zhong, Liming Liu, Xinggong Zhang
Abstract
AI Video Assistant emerges as a new paradigm for Realtime Communication (RTC), where one peer is a Multimodal Large Language Model (MLLM) deployed in the cloud. This makes interaction between humans and AI more intuitive, akin to chatting with a real person. However, a fundamental mismatch exists between current RTC frameworks and AI Video Assistants, stemming from the drastic shift in Quality of Experience (QoE) and more challenging networks. Measurements on our production prototype also confirm that current RTC fails, causing latency spikes and accuracy drops.
To address these challenges, we propose Artic, an AIoriented RTC framework for MLLM Video Assistants, exploring the shift from "humans watching video" to "AI understanding video. " Specifically, Artic proposes: (1) Response Capability-aware Adaptive Bitrate, which utilizes MLLM accuracy saturation to proactively cap bitrate, reserving bandwidth headroom to absorb future fluctuations for latency reduction; (2) Zero-overhead Context-aware Streaming, which allocates limited bitrate to regions most important for the response, maintaining accuracy even under ultra-low bitrates; and (3) Degraded Video Understanding Benchmark, the first benchmark evaluating how RTC-induced video degradation affects MLLM accuracy. Prototype experiments using realworld uplink traces show that compared with existing methods, Artic significantly improves accuracy by 15.12% and reduces latency by 135.31 ms. We will release the benchmark and codes at https://github.com/pku-netvideo/DeViBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c48e5d17-00fa-4af1-8868-dc0c31589d16Builds on14
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang et al.NeurIPS 2025 · 234 citations
- OnRL: improving mobile video telephony via online reinforcement learningHuanhuan Zhang, Anfu Zhou, Jiamin Lu, Ruoxuan Ma et al.MobiCom 2020 · 105 citations
- Loki: improving long tail performance of learning-based real-time video adaptation by fusing rule-based modelsHuanhuan Zhang, Anfu Zhou, Yuhan Hu, Chaoyue Li et al.MobiCom 2021 · 78 citations
- Tambur: Efficient loss recovery for videoconferencing via streaming codesMichael Rudow, Francis Y. Yan, Abhishek Kumar, Ganesh Ananthanarayanan et al.NSDI 2023 · 77 citations
Related papers
- Proact-VL: A Proactive VideoLLM for Real-Time AI CompanionsWeicai Yan, Yuhong Dai, Qi Ran, Haodong Li et al.ICML 2026 · 6 citations
- RIVER: A Real-Time Interaction Benchmark for Video LLMsYansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng et al.ICLR 2026 · 12 citations
- UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLMKaijie Xiao, Yi Gao, Wei DongUbiComp 2026
- LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life TasksHengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen et al.CVPR 2026 · 4 citations
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge et al.ICLR 2025
