RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
Chenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao, Chunliang Zhang, Tongran Liu, Jingbo Zhu
Abstract
Large vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd7114c5-5a27-4478-a9c9-ddad3ebebddcCited by top-tier papers5
- VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin et al.ACL 2026 · 9 citations
- SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementYuan Ge, Junxiang Zhang, Xiaoqian Liu, Bei Li et al.AAAI 2026 · 5 citations
- MSRL: Scaling Generative Multimodal Reward Modeling via Multi-Stage Reinforcement LearningChenglong Wang, Yifu Huo, Yang Gan, Qiaozhi He et al.CVPR 2026 · 5 citations
- GRAM: A Generative Foundation Reward Model for Reward GeneralizationChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu et al.ICML 2025
- SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationJianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye et al.CVPR 2025
Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao et al.NeurIPS 2023 · 810 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
Related papers
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackTianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He et al.CVPR 2024 · 72 citations
- Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference OptimizationShuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai et al.EMNLP 2025 · 2 citations
- MM-RLHF: The Next Step Forward in Multimodal LLM AlignmentYifan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu et al.ICML 2025
- Calibrated Self-Rewarding Vision Language ModelsYiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang et al.NeurIPS 2024 · 77 citations
- Self-Supervised Visual Preference AlignmentKe Zhu, Liang Zhao, Zheng Ge, Xiangyu ZhangACM MM 2024 · 7 citations
