MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Yifan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu, Peiyan Li, Jianshu Zeng, Wulin Xie, Yang Shi, Huanyu Zhang, Junkang Wu, Xue Wang, Yibo Hu
Abstract
Despite notable advancements in Multimodal Large Language Models (MLLMs), most state-of-the-art models have not undergone thorough alignment with human preferences. This gap exists because current alignment research has primarily achieved progress in specific areas (e.g., hallucination reduction), while the broader question of whether aligning models with human preferences can systematically enhance MLLM capability remains largely unexplored. To this end, we introduce MM-RLHF, a dataset containing 120k fine-grained, human-annotated preference comparison pairs. This dataset represents a substantial advancement over existing resources, offering superior size, diversity, annotation granularity, and quality. Leveraging this dataset, we propose several key innovations to improve both the quality of reward models and the efficiency of alignment algorithms. Notably, we introduce a Critique-Based Reward Model, which generates critiques of model outputs before assigning scores, offering enhanced interpretability and more informative feedback compared to traditional scalar reward mechanisms. Additionally, we propose Dynamic Reward Scaling, a method that adjusts the loss weight of each sample according to the reward signal, thereby optimizing the use of high-quality comparison pairs. Our approach is rigorously evaluated across 10 distinct dimensions and 27 benchmarks, with results demonstrating significant and consistent improvements in performance. Specifically, fine-tuning LLaVA-ov-7B with MM-RLHF and our alignment algorithm leads to a 19.5% increase in conversational abilities and a 60% improvement in safety.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea8d80ab-050e-42d1-b3b2-408a9710e5d1Cited by top-tier papers35
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- Thyme: Think Beyond ImagesYifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu et al.ICLR 2026 · 146 citations
- NoisyRollout: Reinforcing Visual Reasoning with Data AugmentationXiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du et al.NeurIPS 2025 · 104 citations
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu et al.ICLR 2026 · 65 citations
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation ModelsWulin Xie, YiFan Zhang, Chaoyou Fu, Yang Shi et al.ICLR 2026 · 31 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
Related papers
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackTianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He et al.CVPR 2024 · 72 citations
- GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward ModelJingyu Zhang, Kun Yang, Ming Wen, jiawei zhao et al.ICML 2026
- LLaVA-Critic: Learning to Evaluate Multimodal ModelsTianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye et al.CVPR 2025
- RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataChenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu et al.AAAI 2025
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackYafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li et al.ICML 2025
