R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Yifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu, Bin Wen, Tianke Zhang, Changyi Liu, Kaiyu Jiang, Kaibing Chen, Kaiyu Tang, Haojie Ding, Jiankang Chen
Abstract
Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on improving the model structure and training data of MRMs, there has been limited exploration into the effectiveness of long-term reasoning capabilities for reward modeling and how to activate these capabilities in MRMs. In this paper, we explore how Reinforcement Learning (RL) can be used to improve reward modeling. Specifically, we reformulate the reward modeling problem as a rule-based RL task. However, we observe that directly applying existing RL algorithms, such as Reinforce++, to reward modeling often leads to training instability or even collapse due to the inherent limitations of these algorithms. To address this issue, we propose the StableReinforce algorithm, which refines the training loss, advantage estimation strategy, and reward design of existing RL methods. These refinements result in more stable training dynamics and superior performance. To facilitate MRM training, we collect 200K preference data from diverse datasets. Our reward model, R1-Reward, trained using the StableReinforce algorithm on this dataset, significantly improves performance on multimodal reward modeling benchmarks. Compared to previous SOTA models, R1-Reward achieves a improvement on the VL Reward-Bench and a improvement on the Multimodal Reward Bench. Moreover, with more inference compute, R1-Reward's performance is further enhanced, highlighting the potential of RL algorithms in optimizing MRMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8c30652-9b9c-4b74-90fa-ea5264334d71Cited by top-tier papers18
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu et al.ICLR 2026 · 146 citations
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu et al.NeurIPS 2025 · 33 citations
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsYunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang et al.NeurIPS 2025 · 18 citations
- BaseReward: A Strong Baseline for Multimodal Reward ModelYiFan Zhang, Haihua Yang, Huanyu Zhang, Yang Shi et al.ICLR 2026 · 16 citations
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang et al.CVPR 2026 · 16 citations
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang et al.NeurIPS 2025 · 234 citations
Related papers
- Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement LearningXiaodong Wang, Zhirong Wu, Langling Huang, Yuxi Zheng et al.CVPR 2026
- MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement LearningChenglong Wang, Yifu Huo, Yang Gan, Qiaozhi He et al.CVPR 2026 · 5 citations
- The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsZichao Li, Xueru Wen, Jie Lou, Yuqiu Ji et al.ICML 2025
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun et al.NeurIPS 2025 · 17 citations
- Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement LearningShuang Chen, Hangyu Guo, Zhaochen Su, Yafu Li et al.ICLR 2026 · 49 citations
