Lune

NeurIPS2025顶会

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Romy Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

2025年份
21被引次数
5顶会引用

摘要

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper presents a systematic analysis revealing that CoT often degrades performance in video reasoning, generating verbose but misleading internal monologues, and leading to hallucinated visual details and overridden correct intuitions-a phenomenon we term "visual thinking drift." We explain this drift through a Bayesian lens, positing that CoT traces often diverge from actual visual evidence, instead amplifying internal biases or language priors, causing models to storytell rather than engage in grounded reasoning. To counteract this, we introduce Visual Evidence Reward (VER), a reinforcement learning framework that explicitly rewards the generation of reasoning traces that are verifiably grounded in visual evidence. Comprehensive evaluation across 10 diverse video understanding benchmarks demonstrates that our Video-VER consistently achieves top performance. Our work sheds light on the distinct challenges of video-centric reasoning and encourages the development of AI that robustly grounds its inferences in visual evidence-for large multimodal models that not only "think before answering", but also "see while thinking". 1 Next, to counter the visual thinking drift problem, we introduce Visual Evidence Reward (VER), a novel reward mechanism for reinforcement learning-based MLLM post-training framework. Our VER explicitly encourages reasoning traces grounded in visual evidence. The key insight is that genuine video reasoning emerges when the internal thought process itself is actively and granularly tethered to perceived content, compelling models to truly "see while thinking", not just "think before answering". In the proposed model, an auxiliary LLM acts as a judge, evaluating the factual alignment between intermediate thoughts and visual inputs. This automatically encourages not just coherent but correct reasoning, stabilizing inference, and boosting overall accuracy.

We evaluate our Video-VER model across 10 diverse video understanding benchmarks. Compared to strong base models and existing reasoning techniques, Video-VER consistently ranks first or second. Furthermore, our model achieves consistently strong margins compared to its respective base MLLM (trained without the Visual Evidence Reward)-as much as +9.0% absolute accuracy gains, and an average of +4.0% across all 10 benchmarks. Our results suggest that in video reasoning, grounding-not verbosity-is essential to true video intelligence.

2 Related Work Eliciting Reasoning Ability from Large Language Models Large language models (LLMs) have shown strong performance on complex reasoning tasks such as mathematics and programming [7,11,50,2,47]. These capabilities are often elicited through few-shot [58,8,64,78] and zeroshot prompting [65,27], or through instruction tuning with large-scale chain-of-thought (CoT)

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper30

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖