GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model
Jingyu Zhang, Kun Yang, Ming Wen, jiawei zhao, Yuxuan Liu, Zhuoer Xu, shiwen cui
摘要
Aligning Multimodal Large Language Models (MLLMs) with human preferences remains a fundamental challenge. While Generative Reward Models (GRMs) offer a promising reasoning-based alternative to scalar models, they are often hindered by severe position bias and prohibitively high computational overhead. To address these limitations, we propose GenAlign, a unified framework that synergizes robust generative reward modeling with efficient MLLM alignment. First, we introduce a rubric-based GRM that explicitly models the preference judgment process. By employing reinforcement learning with verifiable rewards and an online position debiasing mechanism, our model produces interpretable reasoning critiques and robust preference predictions. Second, we propose a policy optimization strategy utilizing advantage-smoothed dynamic reference anchoring. This approach reduces computational complexity while mitigating gradient instability caused by variance collapse. Extensive experiments demonstrate that GenAlign achieves state-of-the-art preference prediction accuracy on multimodal reward modeling benchmarks. Moreover, it consistently improves the performance of three MLLMs across seven diverse evaluation benchmarks, particularly making significant progress in safety and hallucination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
相关 Paper
- Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceJiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun 等NeurIPS 2025 · 被引用 17 次
- MM-RLHF: The Next Step Forward in Multimodal LLM AlignmentYifan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu 等ICML 2025
- ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment FrameworkKai Qin, Liangxin Liu, Yu Liang, Longzheng Wang 等ACL 2026
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu 等EMNLP 2024 · 被引用 11 次
- Aligning Deep Implicit Preferences by Learning to Reason DefensivelyPeiming Li, Zhiyuan Hu, Shiyu Li, Xi Chen 等ICLR 2026
