Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning
Xiaodong Wang, Zhirong Wu, Langling Huang, Yuxi Zheng, Peixi Peng
摘要
Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement learning (RL) to boost models' deep linguistic reasoning of complex videos. But these methods have two main problems: First, the RL framework they used has unstable training, high training costs, and is difficult to train satisfactory video reasoning models; Second, the linguistic reasoning process is difficult to guarantee the reliability of visual information. To alleviate these problems, we propose to use multimodal elements for reasoning, and we design a novel framework to build and enhance versatile video reasoning capabilities on MLLMs. We carefully design a multi-task cold start and multi-task reinforcement learning to improve the model's visual perception and proficiency in multiple capabilities. In the inference phase, we leverage multimodal reasoning and dynamic sampling to further improve the performance. We verified the efficiency of the framework on a base MLLM (Qwen2-VL-7B-Base). Through cold-start with 3k data and reinforcement learning training with 5k data, combined with inference design, our final model significantly outperforms the base model on seven public video benchmarks, even surpassing and approaching the state-of-the-art Instruct Models such as Qwen2.5-VL-7B-Instruct trained with large-scale data. Our code will be available at https://github.com/ Wang-Xiaodong1899/VideoReasoner.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho 等NeurIPS 2025 · 被引用 359 次
相关 Paper
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai 等CVPR 2026 · 被引用 17 次
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual ReasoningYana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin 等NeurIPS 2025 · 被引用 39 次
- VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video ReasoningZhongan Wang, Xiaoyu Wen, Lingxiao Du, Kun Li 等CVPR 2026
- Scaling RL to Long VideosYukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu 等NeurIPS 2025 · 被引用 91 次
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma 等CVPR 2026 · 被引用 92 次
