VR-Thinker: Boosting Multimodal Reward Models through Think with Image Reasoning
Qunzhong Wang, Jie Liu, Jiajun Liang, Yuanxing Zhang, Yilei Jiang, Yaozhi ZHENG, Xintao Wang, Pengfei Wan, Xiangyu Yue, Jiaheng Liu
摘要
Recent advancements in multimodal reward models (RMs) have substantially improved posttraining for visual generative models. However, current RMs face inherent limitations: (1) visual inputs consume large context budgets, forcing fewer frames and causing loss of fine-grained details; and (2) all visual information is packed into the initial prompt, exacerbating hallucination and forgetting during chain-of-thought reasoning. To overcome these issues, we introduce VideoReward Thinker a (VR-Thinker), a thinking-with-image framework that equips the RM with visual reasoning operations (e.g., select frame) and a configurable visual memory window. This allows the RM to actively acquire and update visual evidence within context limits, improving reasoning fidelity and reliability. We activate visual reasoning via a reinforcement fine-tuning pipeline: (i) Cold Start with curated visual chain-of-thought data to distill basic reasoning skills and operation formatting; (ii) select samples whose per-dimension and overall judgments are all correct, then conduct Rejection sampling Fine-Tuning on these high-quality traces to further enhance reasoning; and (iii) apply Group Relative Policy Optimization (GRPO) to strengthen reasoning. Our approach delivers state-ofthe-art accuracy among open-source models on video preference benchmarks, especially for longer videos: a 7B VR-Thinker achieves 80.5% on VideoGen Reward, 82.3% on GenAI-Bench, and 75.6% on MJ-Bench-Video. These results validate the effectiveness and promise of thinking-with-image multimodal reward modeling. a https://vr-thinker.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan 等NeurIPS 2025 · 被引用 284 次
- Video-RAG: Visually-aligned Retrieval-Augmented Long Video ComprehensionYongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin 等NeurIPS 2025 · 被引用 164 次
相关 Paper
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame SpotlightingZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang 等ICLR 2026 · 被引用 34 次
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningYibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang 等NeurIPS 2025 · 被引用 102 次
- When Thinking Drifts: Evidential Grounding for Robust Video ReasoningRomy Luo, Zihui Xue, Alex Dimakis, Kristen GraumanNeurIPS 2025 · 被引用 21 次
- VideoTrace-R1: Long Video-based Retrieval-Augmented Generation via Reinforcement LearningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen 等ICML 2026
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-TuningQi (Cheems) Wang, Yanrui Yu, Ye Yuan, Rui Mao 等NeurIPS 2025 · 被引用 103 次
