R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Yifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu, Bin Wen, Tianke Zhang, Changyi Liu, Kaiyu Jiang, Kaibing Chen, Kaiyu Tang, Haojie Ding, Jiankang Chen
摘要
Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on improving the model structure and training data of MRMs, there has been limited exploration into the effectiveness of long-term reasoning capabilities for reward modeling and how to activate these capabilities in MRMs. In this paper, we explore how Reinforcement Learning (RL) can be used to improve reward modeling. Specifically, we reformulate the reward modeling problem as a rule-based RL task. However, we observe that directly applying existing RL algorithms, such as Reinforce++, to reward modeling often leads to training instability or even collapse due to the inherent limitations of these algorithms. To address this issue, we propose the StableReinforce algorithm, which refines the training loss, advantage estimation strategy, and reward design of existing RL methods. These refinements result in more stable training dynamics and superior performance. To facilitate MRM training, we collect 200K preference data from diverse datasets. Our reward model, R1-Reward, trained using the StableReinforce algorithm on this dataset, significantly improves performance on multimodal reward modeling benchmarks. Compared to previous SOTA models, R1-Reward achieves a improvement on the VL Reward-Bench and a improvement on the Multimodal Reward Bench. Moreover, with more inference compute, R1-Reward's performance is further enhanced, highlighting the potential of RL algorithms in optimizing MRMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu 等ICLR 2026 · 被引用 146 次
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu 等NeurIPS 2025 · 被引用 33 次
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsYunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang 等NeurIPS 2025 · 被引用 18 次
- BaseReward: A Strong Baseline for Multimodal Reward ModelYiFan Zhang, Haihua Yang, Huanyu Zhang, Yang Shi 等ICLR 2026 · 被引用 16 次
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang 等CVPR 2026 · 被引用 16 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
相关 Paper
- Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement LearningXiaodong Wang, Zhirong Wu, Langling Huang, Yuxi Zheng 等CVPR 2026
- MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement LearningChenglong Wang, Yifu Huo, Yang Gan, Qiaozhi He 等CVPR 2026 · 被引用 5 次
- The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsZichao Li, Xueru Wen, Jie Lou, Yuqiu Ji 等ICML 2025
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun 等NeurIPS 2025 · 被引用 17 次
- Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement LearningShuang Chen, Hangyu Guo, Zhaochen Su, Yafu Li 等ICLR 2026 · 被引用 49 次
