BaseReward: A Strong Baseline for Multimodal Reward Model
YiFan Zhang, Haihua Yang, Huanyu Zhang, Yang Shi, Zezhou Chen, Haochen Tian, Chaoyou Fu, Kai WU, Bo Cui, Xu Wang, Jianfei Pan, Haotian Wang
摘要
The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently lacking in both academia and industry. Through exhaustive experimental analysis, this paper aims to provide a clear “recipe” for constructing high-performance MRMs. We systematically investigate every crucial component in the MRM development pipeline, including reward modeling paradigms (e.g., Naive-RM, Critic-based RM, and Generative RM), reward head architecture, training strategies, data curation (covering over ten multimodal and text-only preference datasets), backbone model and model scale, and ensemble methods.
Based on these experimental insights, we introduce BaseReward, a powerful and efficient baseline for multimodal reward modeling. BaseReward adopts a simple yet effective architecture, built upon a Qwen2.5-VL backbone, featuring an optimized two-layer reward head, and is trained on a carefully curated mixture of high-quality multimodal and text-only preference data. Our results show that BaseReward establishes a new state-of-the-art (SOTA) on major benchmarks such as MM-RLHF-Reward Bench, VL-Reward Bench, and Multimodal Reward Bench, outperforming previous open-source and proprietary models. Furthermore, to validate its practical utility beyond static benchmarks, we integrate BaseReward into a real-world reinforcement learning pipeline, successfully enhancing an MLLM’s performance across various perception, reasoning, and conversational tasks. This work not only delivers a top-tier MRM but, more importantly, provides the community with a clear, empirically backed guide for developing robust reward models for the next generation of MLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang 等CVPR 2026 · 被引用 16 次
- PhyCritic: Multimodal Critic Models for Physical AITianyi Xiong, Shihao Wang, Guilin Liu, Yi Dong 等CVPR 2026 · 被引用 12 次
- PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design GenerationJianyu LAI, Sixiang Chen, Jialin Gao, Hengyu Shi 等CVPR 2026 · 被引用 5 次
- MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement LearningChenglong Wang, Yifu Huo, Yang Gan, Qiaozhi He 等CVPR 2026 · 被引用 5 次
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward ModelingJianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li 等CVPR 2026
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
相关 Paper
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu 等ICLR 2026 · 被引用 65 次
- The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsZichao Li, Xueru Wen, Jie Lou, Yuqiu Ji 等ICML 2025
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun 等NeurIPS 2025 · 被引用 17 次
- GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward ModelJingyu Zhang, Kun Yang, Ming Wen, jiawei zhao 等ICML 2026
- ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment FrameworkKai Qin, Liangxin Liu, Yu Liang, Longzheng Wang 等ACL 2026
