SPIKE-RL: Video-LLMs meet Bayesian Surprise
Sahithya Ravi, Aditya Chinchure, Raymond T. Ng, Leonid Sigal, Vered Shwartz
Abstract
Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce SPIKE, an inference-time framework that quantifies Bayesian Surprise as the belief update triggered by new visual evidence in the video stream, identifying moments where new visual evidence conflicts with prior beliefs. SPIKE effectively localizes surprise in videos, strongly correlated with humans on positive (FunQA) and negative (Oops!) surprise benchmarks. Since the beliefs of zero-shot Video-LLMs are often suboptimal, we develop SPIKE-RL, which leverages GRPO to optimize belief hypotheses based on a reward signal from the video caption. SPIKE and SPIKE-RL guide query-agnostic surprise-weighted frame sampling, which allocates more frames to interesting moments in the video. With this strategy, we achieve consistent performance gains on five downstream benchmarks over uniform sampling. By enabling Video-LLMs to track beliefs and register surprise, our work paves the way for more robust models that can revise their understanding in response to new information. Code is available at https://github.com/sahithyaravi/SPIKE-RL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fea1bf8a-85f6-4301-b75e-ee883db2f2feBuilds on20
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video UnderstandingWeiyu Guo, Ziyang Chen, Shaoguang Wang, JianXiang He et al.NeurIPS 2025 · 35 citations
Related papers
- SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMsChong Tang, Sannara Ek, Dirk Koch, Robert Mullins et al.ICLR 2026
- MVP: Enhancing Video Large Language Models via Self-supervised Masked Video PredictionXiaokun Sun, Zezhong Wu, Zewen Ding, Linli XuACL 2026 · 1 citation
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language UnderstandingCanhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou et al.AAAI 2026 · 15 citations
- Reinforcing Structured Chain-of-Thought for Video UnderstandingPeiyao Wang, Haotian Xu, Noranart Vesdapunt, Rui Hou et al.CVPR 2026 · 1 citation
- Towards Sparse Video Understanding and ReasoningChenwei Xu, Zhen Ye, Shang Wu, Weijian Li et al.CVPR 2026 · 3 citations
