Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
Jialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu, Yirong Sun, Yunpu Ma, Xiaoyu Shen
Abstract
Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities in Chain-of-Thought (CoT) reasoning. However, existing LVLM reasoning paradigms only begin reasoning after the entire video becomes available, introducing unnecessary latency and diminishing attention to early visual cues in dynamic scenes. Inspired by the human ability to think while watching, we introduce a streaming reasoning paradigm for LVLMs, where reasoning unfolds sequentially with incoming frames and deepens after the full video is observed. We instantiate this paradigm through Think-as-You-See (TaYS), a unified framework that enables LVLMs to reason while watching by integrating streaming CoT generation, stream-constrained training, and stream-parallel inference. Specifically, TaYS employs temporally aligned streaming reasoning units with precise CoT supervision, enforces ordered reasoning via streaming attention masks and positional encodings, and utilizes a parallel KV caches mechanism that decouples input encoding from reasoning generation, ensuring alignment and true concurrency. We evaluate TaYS on the Qwen2.5-VL model family across representative video CoT tasks, including event dynamics analysis, causal reasoning, and thematic understanding. Experimental results show that TaYS achieves superior reasoning performance compared with batch-mode CoT, while reducing pre-reasoning latency to under one second and overall answer delay by more than 50%. These findings demonstrate the effectiveness of the streaming paradigm in enabling real-time, human-like reasoning for LVLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a237c95-125d-4655-ab57-03bf891416a3Cited by top-tier papers1
Ask how each one uses itBuilds on27
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou et al.NeurIPS 2023 · 252 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to CognitionHao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang et al.ICML 2024 · 182 citations
- StreamingVLM: Real-Time Understanding for Infinite Video StreamsRuyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He et al.ICLR 2026 · 95 citations
- Efficient Reasoning with Hidden ThinkingXuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi et al.ICML 2026 · 56 citations
Related papers
- StreamingThinker: Large Language Models Can Think While ReadingJunlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma et al.ICLR 2026 · 17 citations
- ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language ModelsYongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen et al.ACM MM 2025 · 4 citations
- TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal UnderstandingLianyu Hu, Xiaoyu Ma, Zeqin Liao, Yang LiuICML 2026 · 2 citations
- CogStream: Context-guided Streaming Video Question AnsweringZicheng Zhao, Kangyu Wang, Shijie Li, Rui Qian et al.AAAI 2026 · 3 citations
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li et al.ICLR 2026 · 55 citations
