VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Lei Li, Yuancheng Wei, Zhihui Xie, Xuqing Yang, Yifan Song, Peiyi Wang, Chenxin An, Tianyu Liu, Sujian Li, Bill Yuchen Lin, Lingpeng Kong, Qi Liu
Abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench's effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson's r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io . * Core contributors. Answer 1 (Chosen) : There are three sinks and three mirrors in the public restroom. Answer 2 (Rejected) : In the image, there are two sinks and four mirrors equipped with sensor taps alongside a matching set of hand dryers hanging above. Llama-3.2-90B Answer 2 is better because it accurately describes the visual elements in the image, fully addresses all aspects of the question. Claude-3.5-Sonnet Answer 2 provides a more accurate, complete, and informative response to the question while maintaining clarity and relevance. It correctly identifies the number of sinks and mirrors and offers additional details about the restroom that enhance the overall understanding of the space. Qwen2-VL-7B Answer 2 is better, which accurately states that there are two sinks and four mirrors in the image. Query: How many sinks and mirrors are there in the public restroom? Figure 1. An example from our VL-RewardBench asking the visual details in a restroom. Open-source VL-GenRMs (Qwen2-VL-7B [48] and Llama-3.2-90B [13]) and the commercial model (Claude-3.5-Sonnet [3]) all fail to provide accurate judgments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efe8befb-93dd-47bf-a64b-73e30e62c026Cited by top-tier papers34
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningYibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang et al.NeurIPS 2025 · 102 citations
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu et al.ICLR 2026 · 65 citations
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking RewardKaixuan Fan, Kaituo Feng, Haoming Lyu, Dongzhan Zhou et al.ICLR 2026 · 54 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
Related papers
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo et al.ICCV 2025 · 22 citations
- ViLBench: A Suite for Vision-Language Process Reward ModelingHaoqin Tu, Weitao Feng, Hardy Chen, Hui Liu et al.EMNLP 2025 · 1 citation
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and ImageYushi Hu, Reyhane Askari Hemmat, Melissa Hall, Emily Dinan et al.CVPR 2026 · 18 citations
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
- Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian et al.CVPR 2024
