Stable and Efficient Single-Rollout RL for Multimodal Reasoning
Rui Liu, Dian Yu, Lei Ke, Haolin Liu, Yujun Zhou, Zhenwen Liang, Haitao Mi, Pratap Tokekar, Dong Yu
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalent group-based algorithms such as GRPO require multi-rollout sampling for each prompt. While more efficient single-rollout variants have recently been explored in text-only settings, we find that they suffer from severe instability in multimodal contexts, often leading to training collapse. To address this sample efficiency-stability trade-off, we introduce (Multimodal Stabilized Single-Rollout), a group-free RLVR framework that achieves both stable optimization and effective multimodal reasoning performance. MSSR achieves this via an entropy-based advantage-shaping mechanism that adaptively regularizes advantage magnitudes, preventing collapse and maintaining training stability. While such mechanisms have been used in group-based RLVR, we show that in the multimodal single-rollout setting they are not merely beneficial but essential for stability. In in-distribution evaluations, MSSR demonstrates superior rollout sample efficiency, achieving similar validation accuracy with half the training steps. When trained for the same number of steps, MSSR's performance surpasses the group-based baseline and shows consistent generalization improvements across five diverse reasoning-intensive benchmarks. Together, these results demonstrate that MSSR enables stable, sample-efficient, and effective RLVR for complex multimodal reasoning tasks. We will release code and checkpoints upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67e14d88-ffaf-4bea-ab14-6797cff2eb1aCited by top-tier papers2
- CAML: Collaborative Auxiliary Modality Learning for Multi-Agent SystemsRui Liu, Yu Shen, Peng Gao, Pratap Tokekar et al.NeurIPS 2025 · 9 citations
- : A Generalist Value Model for Any Policy at State ZeroYi-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun et al.ICML 2026 · 3 citations
Builds on22
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu et al.NeurIPS 2025 · 356 citations
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai et al.AAAI 2026 · 216 citations
Related papers
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu et al.ICLR 2026 · 51 citations
- Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable RewardsHaechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang et al.ICML 2026
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang et al.CVPR 2026 · 4 citations
- Reinforced Efficient Reasoning via Semantically Diverse ExplorationZiqi Zhao, Zhaochun Ren, Jiahong Zou, Liu Yang et al.ACL 2026 · 5 citations
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang et al.NeurIPS 2025 · 3 citations
