MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
Xiaokun Sun, Zezhong Wu, Zewen Ding, Linli Xu
摘要
Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often lacking explicit supervision for intrinsic temporal coherence and inter-frame correlations. This tendency limits the models'ability to capture intricate dynamics and fine-grained visual causality. To explicitly bridge this gap, we propose a novel post-training objective: Masked Video Prediction (MVP). By requiring the model to reconstruct a masked continuous segment from a set of challenging distractors, MVP forces the model to attend to the sequential logic and temporal context of events. To support scalable training, we introduce a scalable data synthesis pipeline capable of transforming arbitrary video corpora into MVP training samples, and further employ Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties. Comprehensive evaluations demonstrate that MVP enhances video reasoning capabilities by directly reinforcing temporal reasoning and causal understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative PerceptionZiang Yan, Yinan He, Xinhao Li, Zhengrong Yue 等NeurIPS 2025 · 被引用 70 次
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPOJinyoung Park, Jeehye Na, Jinyoung Kim, Hyunwoo J. KimNeurIPS 2025 · 被引用 64 次
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language UnderstandingCanhui Tang, Zifan Han, Hongbo Sun, Sanping Zhou 等AAAI 2026 · 被引用 15 次
相关 Paper
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG 等CVPR 2026 · 被引用 7 次
- Reinforcing Structured Chain-of-Thought for Video UnderstandingPeiyao Wang, Haotian Xu, Noranart Vesdapunt, Rui Hou 等CVPR 2026 · 被引用 1 次
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma 等CVPR 2026 · 被引用 92 次
- VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningLinhan Cao, Wei Sun, Weixia Zhang, Xiangyang Zhu 等AAAI 2026 · 被引用 6 次
- Towards Unified Multimodal Interleaved Generation via Group Relative Policy OptimizationMing Nie, Chunwei Wang, Jianhua Han, Hang Xu 等NeurIPS 2025 · 被引用 7 次
