Lune

ICML2026Top-tier venue

Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with Images

Bo-Wen Yin, Qize Yang, Boyuan Sun, Xihan Wei, Qibin Hou

2026Year

Abstract

The "thinking with images" paradigm encourages multimodal large language models to generate intermediate visual steps-such as cropping, annotation, spatial localization, and sketches-to enhance high-resolution perception and complex reasoning. However, existing multimodal Process Reward Models (PRMs) evaluate only textual reasoning and cannot judge the correctness of these visual steps, creating a key gap when visual reasoning is essential for solving tasks. We propose Discriminative Visual Process Reward Model (DiscPRM), a multimodal PRM that jointly evaluates textual and visual intermediate steps by modeling visual reasoning trajectories, image operations, and text-image consistency. To support this, we build VTReward-100K, a dataset of stepby-step visual reasoning sequences with supervision. Experiments show that using DiscPRM for Best-of-N process supervision substantially improves multimodal reasoning performance on tasks requiring visual intermediate steps, achieving over 5% gains across benchmarks. We further introduce VABench, the first benchmark for evaluating PRMs on visual reasoning error detection. We hope this work can provide foundational support for advancing the emerging direction of visual-textual process reward. Project page: https://DiscPRM.github.io.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 73b0e486-321d-4cad-a9a6-3ba43e9ca259

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines