Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with Images
Bo-Wen Yin, Qize Yang, Boyuan Sun, Xihan Wei, Qibin Hou
摘要
The "thinking with images" paradigm encourages multimodal large language models to generate intermediate visual steps-such as cropping, annotation, spatial localization, and sketches-to enhance high-resolution perception and complex reasoning. However, existing multimodal Process Reward Models (PRMs) evaluate only textual reasoning and cannot judge the correctness of these visual steps, creating a key gap when visual reasoning is essential for solving tasks. We propose Discriminative Visual Process Reward Model (DiscPRM), a multimodal PRM that jointly evaluates textual and visual intermediate steps by modeling visual reasoning trajectories, image operations, and text-image consistency. To support this, we build VTReward-100K, a dataset of stepby-step visual reasoning sequences with supervision. Experiments show that using DiscPRM for Best-of-N process supervision substantially improves multimodal reasoning performance on tasks requiring visual intermediate steps, achieving over 5% gains across benchmarks. We further introduce VABench, the first benchmark for evaluating PRMs on visual reasoning error detection. We hope this work can provide foundational support for advancing the emerging direction of visual-textual process reward. Project page: https://DiscPRM.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 236 次
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang 等NeurIPS 2025 · 被引用 234 次
相关 Paper
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen 等ICLR 2026 · 被引用 110 次
- What, Whether and How? Unveiling Process Reward Models for Thinking with Images ReasoningYujin Zhou, Pengcheng Wen, Jiale Chen, Boqin Yin 等AAAI 2026 · 被引用 2 次
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo 等ICCV 2025 · 被引用 22 次
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained RewardsHonghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang 等NeurIPS 2025 · 被引用 7 次
- ViLBench: A Suite for Vision-Language Process Reward ModelingHaoqin Tu, Weitao Feng, Hardy Chen, Hui Liu 等EMNLP 2025 · 被引用 1 次
