When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
Romy Luo, Zihui Xue, Alex Dimakis, Kristen Grauman
Abstract
Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper presents a systematic analysis revealing that CoT often degrades performance in video reasoning, generating verbose but misleading internal monologues, and leading to hallucinated visual details and overridden correct intuitions-a phenomenon we term "visual thinking drift." We explain this drift through a Bayesian lens, positing that CoT traces often diverge from actual visual evidence, instead amplifying internal biases or language priors, causing models to storytell rather than engage in grounded reasoning. To counteract this, we introduce Visual Evidence Reward (VER), a reinforcement learning framework that explicitly rewards the generation of reasoning traces that are verifiably grounded in visual evidence. Comprehensive evaluation across 10 diverse video understanding benchmarks demonstrates that our Video-VER consistently achieves top performance. Our work sheds light on the distinct challenges of video-centric reasoning and encourages the development of AI that robustly grounds its inferences in visual evidence-for large multimodal models that not only "think before answering", but also "see while thinking". 1 Next, to counter the visual thinking drift problem, we introduce Visual Evidence Reward (VER), a novel reward mechanism for reinforcement learning-based MLLM post-training framework. Our VER explicitly encourages reasoning traces grounded in visual evidence. The key insight is that genuine video reasoning emerges when the internal thought process itself is actively and granularly tethered to perceived content, compelling models to truly "see while thinking", not just "think before answering". In the proposed model, an auxiliary LLM acts as a judge, evaluating the factual alignment between intermediate thoughts and visual inputs. This automatically encourages not just coherent but correct reasoning, stabilizing inference, and boosting overall accuracy.
We evaluate our Video-VER model across 10 diverse video understanding benchmarks. Compared to strong base models and existing reasoning techniques, Video-VER consistently ranks first or second. Furthermore, our model achieves consistently strong margins compared to its respective base MLLM (trained without the Visual Evidence Reward)-as much as +9.0% absolute accuracy gains, and an average of +4.0% across all 10 benchmarks. Our results suggest that in video reasoning, grounding-not verbosity-is essential to true video intelligence.
2 Related Work Eliciting Reasoning Ability from Large Language Models Large language models (LLMs) have shown strong performance on complex reasoning tasks such as mathematics and programming [7,11,50,2,47]. These capabilities are often elicited through few-shot [58,8,64,78] and zeroshot prompting [65,27], or through instruction tuning with large-scale chain-of-thought (CoT)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2f85cd7-41b5-4120-809d-8ae558ba6e62Cited by top-tier papers5
- Seeing the Arrow of Time in Large Multimodal ModelsZihui Xue, Romy Luo, Kristen GraumanNeurIPS 2025 · 30 citations
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language ModelsJialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu et al.CVPR 2026 · 6 citations
- VideoSSR: Video Self-Supervised Reinforcement LearningZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang et al.CVPR 2026 · 4 citations
- Reinforcing Structured Chain-of-Thought for Video UnderstandingPeiyao Wang, Haotian Xu, Noranart Vesdapunt, Rui Hou et al.CVPR 2026 · 1 citation
- Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video UnderstandingHoulun Chen, Xin Wang, Guangyao Li, Yuwei Zhou et al.SIGIR 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
Related papers
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng et al.ICLR 2026 · 24 citations
- ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language ModelsYongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen et al.ACM MM 2025 · 4 citations
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsYongchang Zhang, Xianzheng Ma, Tianyi Liu, Guangquan Zhou et al.CVPR 2026 · 2 citations
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen et al.CVPR 2026 · 18 citations
