Training Data Efficiency in Multimodal Process Reward Models
Jinyuan Li, Chengsong Huang, Langlin Huang, Shaoyang Xu, Haolin Liu, Wenxuan Zhang, Jiaxin Huang
Abstract
Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training. Our preliminary experiments reveal that MPRM training quickly saturates under random subsampling of the training data, indicating substantial redundancy within existing MC-annotated corpora. To explain this, we formalize a theoretical framework and reveal that informative gradient updates depend on two factors: label mixtures of positive/negative steps and label reliability (average MC scores of positive steps). Guided by these insights, we propose the Balanced-Information Score (BIS), which prioritizes both mixture and reliability based on existing MC signals at the rollout level, without incurring any additional cost. Across two backbones (InternVL2.5-8B and Qwen2.5-VL-7B) on VisualProcessBench, BIS-selected subsets consistently match and even surpass the full-data performance at small fractions. Notably, the BIS subset reaches full-data performance using only 10% of the training data, improving over random subsampling by a relative 4.1%. Our code is released Balanced-Info-MPRM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f13ef4b0-5861-41de-8ad6-a3b1107e1cb1Builds on16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- AlphaMath Almost Zero: Process Supervision without ProcessGuoxin Chen, Minpeng Liao, Chengxi Li, Kai FanNeurIPS 2024 · 219 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
Related papers
- VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward ModelsWeiyun Wang, Zhangwei Gao, Lianjie Chen, Zhe Chen et al.ICLR 2026 · 110 citations
- Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level OptimizationShiping Gao, Hongzhan Chen, Xiaojun Quan, Qifan Wang et al.ICML 2026
- Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with ImagesBo-Wen Yin, Qize Yang, Boyuan Sun, Xihan Wei et al.ICML 2026
- DreamPRM: Domain-reweighted Process Reward Model for Multimodal ReasoningQi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula et al.NeurIPS 2025 · 17 citations
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward LearningYuyang Ding, Xinyu Shi, Juntao Li, Xiaobo Liang et al.NeurIPS 2025 · 10 citations
