What, Whether and How? Unveiling Process Reward Models for Thinking with Images Reasoning
Yujin Zhou, Pengcheng Wen, Jiale Chen, Boqin Yin, Han Zhu, Jiaming Ji, Juntao Dai, Chi-Min Chan, Sirui Han
摘要
The rapid advancement of Large Vision Language Models (LVLMs) has demonstrated excellent abilities in various visual tasks. Building upon these developments, the thinking with images paradigm has emerged, enabling models to dynamically edit and re-encode visual information at each reasoning step, mirroring human visual processing. However, this paradigm introduces significant challenges as diverse errors may occur during reasoning processes. This necessitates Process Reward Models (PRMs) for distinguishing positive and negative reasoning steps, yet existing benchmarks for PRMs are predominantly text-centric and lack comprehensive assessment under this paradigm. To address these gaps, this work introduces the first comprehensive benchmark specifically designed for evaluating PRMs under the thinking with images paradigm. Our main contributions are: (1) Through extensive analysis of reasoning trajectories and guided search experiments with PRMs, we define 7 fine-grained error types and demonstrate both the necessity for specialized PRMs and the potential for improvement. (2) We construct a comprehensive benchmark comprising 1,206 manually annotated thinking with images reasoning trajectories spanning 4 categories and 16 subcategories for fine-grained evaluation of PRMs. (3) Our experimental analysis reveals that current LVLMs fall short as effective PRMs, exhibiting limited capabilities in visual reasoning process evaluation with significant performance disparities across error types, positive evaluation bias, and sensitivity to reasoning step positions. These findings demonstrate the effectiveness of our benchmark and establish crucial foundations for advancing PRMs in LVLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic VerificationChuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan 等ICML 2026 · 被引用 3 次
- Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin 等ACL 2026
- Benchmarking Fine-Grained Error Detection in Multimodal ReasoningChi-Min Chan, Han Zhu, Chunyang Jiang, Jiaming Ji 等ACL 2026
它引用的顶会 Paper15
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth 等NeurIPS 2024 · 被引用 373 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 等NeurIPS 2025 · 被引用 153 次
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain 等NeurIPS 2025 · 被引用 90 次
相关 Paper
- Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with ImagesBo-Wen Yin, Qize Yang, Boyuan Sun, Xihan Wei 等ICML 2026
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo 等ICCV 2025 · 被引用 22 次
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen 等ICLR 2026 · 被引用 110 次
- ViLBench: A Suite for Vision-Language Process Reward ModelingHaoqin Tu, Weitao Feng, Hardy Chen, Hui Liu 等EMNLP 2025 · 被引用 1 次
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained RewardsHonghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang 等NeurIPS 2025 · 被引用 7 次
