Lune

NeurIPS2025顶会

DreamPRM: Domain-reweighted Process Reward Model for Multimodal Reasoning

Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

2025年份
17被引次数
1顶会引用

摘要

Reasoning has substantially improved the performance of large language models (LLMs) on complicated tasks. Central to the current reasoning studies, Process Reward Models (PRMs) offer a fine-grained evaluation of intermediate reasoning steps and guide the reasoning process. However, extending PRMs to multimodal large language models (MLLMs) introduces challenges. Since multimodal reasoning covers a wider range of tasks compared to text-only scenarios, the resulting distribution shift from the training to testing sets is more severe, leading to greater generalization difficulty. Training a reliable multimodal PRM, therefore, demands large and diverse datasets to ensure sufficient coverage. However, current multimodal reasoning datasets suffer from a marked quality imbalance, which degrades PRM performance and highlights the need for an effective data selection strategy. To address the issues, we introduce DreamPRM, a domain-reweighted training framework for multimodal PRMs which employs bi-level optimization. In the lower-level optimization, DreamPRM performs fine-tuning on multiple datasets with domain weights, allowing the PRM to prioritize high-quality reasoning signals and alleviating the impact of dataset quality imbalance. In the upper-level optimization, the PRM is evaluated on a separate meta-learning dataset; this feedback updates the domain weights through an aggregation loss function, thereby improving the generalization capability of trained PRM. Extensive experiments on multiple multimodal reasoning benchmarks covering both mathematical and general reasoning show that test-time scaling with DreamPRM consistently improves the performance of state-of-the-art MLLMs. Further comparisons reveal that DreamPRM's domain-reweighting strategy surpasses other data selection methods and yields higher accuracy gains than existing test-time scaling approaches. Notably, DreamPRM achieves a top-1 accuracy of 85.2% on the MATHVISTA leaderboard using the o4-mini model, demonstrating its strong generalization in complex multimodal reasoning tasks. 39th Conference on Neural Information Processing Systems (NeurIPS 2025). Dataset difficulty: easy (InternVL-2.5-MPO-8B's accuracy 84.6%) Unnecessary modality: can answer without image Requirements for reasoning: do not require complicated reasoning Domain weight: 0.55 (Determined by DreamPRM) Question: What does the bird feed on? Choices: A. zooplankton B. grass C. predator fish D. none of the above Answer: C Dataset: AI2D (2016) Dataset difficulty: hard (InternVL-2.5-MPO-8B's accuracy 62.1%) Unnecessary modality: cannot answer without image Requirements for reasoning ability: require complicated reasoning Domain weight: 1.49 (Determined by DreamPRM) Question: Determine the scientific nomenclature of the organism shown in the primary image. Choices: A. Hemidactylus turcicus B. Felis silvestris C. Macropus agilis D. None of the above Answer: D Dataset: M3CoT (2024) Figure 1: DreamPRM improves multimodal reasoning by mitigating the dataset quality imbalance problem. Left: On five benchmarks, DreamPRM outperforms base model (InternVL-2.5-8B-MPO [67]) by an average of +4.0%. DreamPRM also consistently surpasses Vanilla PRM trained without data selection. Right: Easy AI2D [23] questions (weight 0.55) vs. hard M3COT [6] questions (weight 1.49) shows how DreamPRM prioritizes data that demand deeper reasoning -samples requiring knowledge from both textual and visual modalities for step-by-step logical deduction.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖