Lune

NeurIPS2025Top-tier venue

DreamPRM: Domain-reweighted Process Reward Model for Multimodal Reasoning

Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

2025Year
17Citations
1Top-tier citations

Abstract

Reasoning has substantially improved the performance of large language models (LLMs) on complicated tasks. Central to the current reasoning studies, Process Reward Models (PRMs) offer a fine-grained evaluation of intermediate reasoning steps and guide the reasoning process. However, extending PRMs to multimodal large language models (MLLMs) introduces challenges. Since multimodal reasoning covers a wider range of tasks compared to text-only scenarios, the resulting distribution shift from the training to testing sets is more severe, leading to greater generalization difficulty. Training a reliable multimodal PRM, therefore, demands large and diverse datasets to ensure sufficient coverage. However, current multimodal reasoning datasets suffer from a marked quality imbalance, which degrades PRM performance and highlights the need for an effective data selection strategy. To address the issues, we introduce DreamPRM, a domain-reweighted training framework for multimodal PRMs which employs bi-level optimization. In the lower-level optimization, DreamPRM performs fine-tuning on multiple datasets with domain weights, allowing the PRM to prioritize high-quality reasoning signals and alleviating the impact of dataset quality imbalance. In the upper-level optimization, the PRM is evaluated on a separate meta-learning dataset; this feedback updates the domain weights through an aggregation loss function, thereby improving the generalization capability of trained PRM. Extensive experiments on multiple multimodal reasoning benchmarks covering both mathematical and general reasoning show that test-time scaling with DreamPRM consistently improves the performance of state-of-the-art MLLMs. Further comparisons reveal that DreamPRM's domain-reweighting strategy surpasses other data selection methods and yields higher accuracy gains than existing test-time scaling approaches. Notably, DreamPRM achieves a top-1 accuracy of 85.2% on the MATHVISTA leaderboard using the o4-mini model, demonstrating its strong generalization in complex multimodal reasoning tasks. 39th Conference on Neural Information Processing Systems (NeurIPS 2025). Dataset difficulty: easy (InternVL-2.5-MPO-8B's accuracy 84.6%) Unnecessary modality: can answer without image Requirements for reasoning: do not require complicated reasoning Domain weight: 0.55 (Determined by DreamPRM) Question: What does the bird feed on? Choices: A. zooplankton B. grass C. predator fish D. none of the above Answer: C Dataset: AI2D (2016) Dataset difficulty: hard (InternVL-2.5-MPO-8B's accuracy 62.1%) Unnecessary modality: cannot answer without image Requirements for reasoning ability: require complicated reasoning Domain weight: 1.49 (Determined by DreamPRM) Question: Determine the scientific nomenclature of the organism shown in the primary image. Choices: A. Hemidactylus turcicus B. Felis silvestris C. Macropus agilis D. None of the above Answer: D Dataset: M3CoT (2024) Figure 1: DreamPRM improves multimodal reasoning by mitigating the dataset quality imbalance problem. Left: On five benchmarks, DreamPRM outperforms base model (InternVL-2.5-8B-MPO [67]) by an average of +4.0%. DreamPRM also consistently surpasses Vanilla PRM trained without data selection. Right: Easy AI2D [23] questions (weight 0.55) vs. hard M3COT [6] questions (weight 1.49) shows how DreamPRM prioritizes data that demand deeper reasoning -samples requiring knowledge from both textual and visual modalities for step-by-step logical deduction.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 708b1f3f-b452-4ae5-ab35-d8d3f67c82f1

Cited by top-tier papers1

Ask how each one uses it

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines