TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
Linli Yao, Yicheng Li, Yuancheng Wei, Lei Li, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Shicheng Li, Sida Li, Lingpeng Kong, Qi Liu
Abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3d15671-7b17-4c4d-807b-30b9ff4d15caCited by top-tier papers22
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li et al.NeurIPS 2025 · 79 citations
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantHaibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu et al.NeurIPS 2025 · 63 citations
- StreamingTOM: Streaming Token Compression for Efficient Video UnderstandingXueyi Chen, Keda Tao, Kele Shao, Huan WangCVPR 2026 · 46 citations
- Accelerating Streaming Video Large Language Models via Hierarchical Token CompressionYiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin et al.CVPR 2026 · 40 citations
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao et al.ICLR 2026 · 37 citations
Builds on27
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng et al.ICLR 2026 · 172 citations
- Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsHongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun et al.CVPR 2022 · 107 citations
Related papers
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge et al.ICLR 2025
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Online Video Understanding: OVBench and VideoChat-OnlineZhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang et al.CVPR 2025
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMsEvangelos Dorovatas, Soroush Seifi, Gunshi Gupta, Rahaf AljundiNeurIPS 2025 · 8 citations
- RIVER: A Real-Time Interaction Benchmark for Video LLMsYansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng et al.ICLR 2026 · 12 citations
