Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
Jiahao Wang, Weiye Xu, Aijun Yang, Wengang Zhou, Lewei Lu, Houqiang Li, Xiaohua Wang, Jinguo Zhu
摘要
Outcome-reward reinforcement learning (RL) is a common-and increasingly significant-way to refine the step-by-step reasoning of multimodal large language models (MLLMs). In the multiple-choice setting-a dominant format for multimodal reasoning benchmarks-the paradigm faces a significant yet often overlooked obstacle: unfaithful trajectories that guess the correct option after a faulty chain of thought receive the same reward as genuine reasoning, which is a flaw that cannot be ignored. We propose Self-Consistency Sampling (SCS) to correct this issue. For each question, SCS (i) introduces small visual perturbations and (ii) performs repeated truncation-and-resampling of an initial trajectory; agreement among the resulting trajectories yields a differentiable consistency score that down-weights unreliable traces during policy updates. Based on Qwen2.5-VL-7B-Instruct, plugging SCS into RLOO, GRPO, and REINFORCE++ series improves accuracy by up to 7.7 percentage points on six multimodal benchmarks with negligible extra computation. SCS also yields notable gains on both Qwen2.5-VL-3B-Instruct and InternVL3-8B, offering a simple, general remedy for outcome-reward RL in MLLMs. Our code is available at https://github.com/GenuineWWD/SCS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu 等NeurIPS 2025 · 被引用 356 次
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 等ACL 2025 · 被引用 236 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
相关 Paper
- First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-TrainingLai Wei, Yuting Li, Chen Wang, Yue Wang 等NeurIPS 2025 · 被引用 28 次
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang 等NeurIPS 2025 · 被引用 45 次
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan 等NeurIPS 2024 · 被引用 27 次
- MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningZhiheng Xi, Yuhui Wang, Yiwen Ding, Guanyu Li 等AAAI 2026
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal ReasoningYongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han 等CVPR 2026
