Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Kecheng Zhang, Zongxin Yang, Mingfei Han, Haihong Hao, Yunzhi Zhuge, Changlin Li, Junhan Zhao, Zhihui Li, Xiaojun Chang
Abstract
Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challenges: a lack of decision transparency, the difficulty of aligning response timing with visual evidence, and the need to maintain a global, causally consistent understanding under tight computational budgets. To address these issues, we propose a novel framework that decouples reasoning control from memory integration. We introduce ****, an instantiation of this framework with two core components. First, the Active Thinking Decision Maker (ATDM) is a transparent reasoning controller that externalizes its decision process using observable progress () and confidence () metrics. This allows it to precisely time its response to match the first-sufficient-evidence timestamp while streaming its reasoning to the user. Second, the Hierarchical Progressive Semantic Integration (HPSI) module acts as an efficient memory system. It employs a set of learnable, multi-level aggregation tokens that are propagated across clips to build a rich, global cognitive state without exceeding token budgets. %Our approach sets a new standard on key online video understanding benchmarks, achieving strong performance of 71.6% on StreamingBench and 46.9% on OVOBench, demonstrating a robust solution for evidence-aligned and transparent online video analysis. Extensive experiments demonstrate the effectiveness of ATDM and HPSI, e.g., Thinking-QwenVL improves the accuracy of the previous state-of-the-art from 67.63% to 71.60% on the StreamingBench benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin et al.NeurIPS 2025 · 164 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng et al.ICCV 2025 · 28 citations
Related papers
- OASIS: On-Demand Hierarchical Event Memory for Streaming Video ReasoningZhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang et al.CVPR 2026 · 16 citations
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video UnderstandingYufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding et al.CVPR 2026 · 24 citations
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?Junbo Niu, Yifei Li, Ziyang Miao, Chunjiang Ge et al.CVPR 2025
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language ModelsJialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu et al.CVPR 2026 · 6 citations
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeHaomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge et al.ICLR 2025
