VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo, Daoxin Zhang, Zhe Xu, Yao Hu, Ting Liu, Yuzhuo Fu
Abstract
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Recently, reward models (RMs) have become increasingly pivotal in the reasoning process. Specifically, process RMs evaluate each reasoning step, outcome RMs focus on the assessment of reasoning results, and critique RMs perform error analysis on the entire reasoning process, followed by corrections. However, existing benchmarks for vision-language RMs (VLRMs) typically assess only a single aspect of their capabilities (e.g., distinguishing between two answers), thus limiting the all-round evaluation and restricting the development of RMs in the visual-language domain. To address this gap, we propose a comprehensive and challenging benchmark, dubbed as VLRMBench, encompassing 12,634 questions. VLRMBench is constructed based on three distinct types of datasets, covering mathematical reasoning, hallucination understanding, and multiimage understanding. We design 12 tasks across three major categories, focusing on evaluating VLRMs in the aspects of process understanding, outcome judgment, and critique generation. Extensive experiments are conducted on 21 open-source models and 5 advanced closed-source models, highlighting the challenges posed by VLRMBench. For instance, in the 'Forecasting Future', a binary classification task, the advanced GPT-4o achieves only a 76.0% accuracy. Additionally, we perform comprehensive analytical studies, offering valuable insights for the future development of VL-RMs. We anticipate that VLRMBench will serve as a pivotal benchmark in advancing VLRMs. Code and datasets will be available at https://github.com/JCruan519/VLRMBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4a0ccb7-d8f6-4afe-a397-dcebd2b92760Cited by top-tier papers6
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form PreferencesZhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li et al.ICLR 2026 · 11 citations
- MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language ModelsYang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu et al.ACL 2026 · 9 citations
- MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-JudgeSua Lee, Sanghee Park, Jinbae ImACL 2026 · 1 citation
- Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin et al.ACL 2026
Builds on22
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji et al.ICML 2024 · 527 citations
Related papers
- VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsLei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang et al.CVPR 2025
- ViLBench: A Suite for Vision-Language Process Reward ModelingHaoqin Tu, Weitao Feng, Hardy Chen, Hui Liu et al.EMNLP 2025 · 1 citation
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and ImageYushi Hu, Reyhane Askari Hemmat, Melissa Hall, Emily Dinan et al.CVPR 2026 · 18 citations
- Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with ImagesBo-Wen Yin, Qize Yang, Boyuan Sun, Xihan Wei et al.ICML 2026
