NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
Longtian Qiu, Shan Ning, Jiaxuan Sun, Xuming He
Abstract
Reinforcement learning (RL) has shown promise in enhancing the general Chainof-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distribution. To address this, we propose NoisyGRPO, a systematic multimodal RL framework that introduces controllable noise into visual inputs for enhanced exploration and explicitly models the advantage estimation process via a Bayesian framework. Specifically, NoisyGRPO improves RL training by: (1) Noise-Injected Exploration Policy: Perturbing visual inputs with Gaussian noise to encourage exploration across a wider range of visual scenarios; and (2) Bayesian Advantage Estimation: Formulating advantage estimation as a principled Bayesian inference problem, where the injected noise level serves as a prior and the observed trajectory reward as the likelihood. This Bayesian modeling fuses both sources of information to compute a robust posterior estimate of trajectory advantage, effectively guiding MLLMs to prefer visually grounded trajectories over noisy ones. Experiments on standard CoT quality, general capability, and hallucination benchmarks demonstrate that NoisyGRPO substantially improves generalization and robustness, especially in RL settings with small-scale MLLMs such as Qwen2.5-VL 3B. The project page is available at https://artanic30.github.io/project_pages/NoisyGRPO.
2 Related Works Multi-modal Large Language Model Vision-language models (VLMs) have rapidly advanced in their ability to comprehend and reason over both visual and textual modalities and demonstrate great improvements over downstream tasks [35,64,12,9,31,40]. These models typically integrate visual encoders with large language models (LLMs) to enable cross-modal understanding and inference. Foundational works such as Flamingo [2] and BLIP-2 [20] established effective strategies for aligning vision and language components. The LLaVA-series [25,24,18] and SPHINXseries [22,10] further advanced the field by introducing visual instruction tuning, significantly enhancing multimodal capabilities. Large-scale models like GPT-4o [14] and Gemini [42] have demonstrated strong general-purpose visual understanding through large-scale multimodal pertaining. To address scalability, mixture-of-experts approaches-such as DeepSeek-VL2 [55], Uni-MoE [21], and MoVA [65]-improve computational efficiency by selectively activating expert modules based on input characteristics. More recently, domain-specific methods such as Math-LLaVA [47] and MAVIS [60] employed mathematical visual instruction tuning to improve VLMs' ability to interpret and solve complex multimodal math problems. At the same time, unified architectures like SEED-X [11], Chameleon [48], Show-O [56], and the Janus series [54, 29, 6] integrate both visual understanding and generation capabilities within a single framework. Despite these advances, most existing VLMs still struggle with robust visual reasoning, particularly in tasks that demand deep visual analysis and complex multi-step reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52bffe01-13b1-4bbf-830c-4d77352bd0cfCited by top-tier papers2
- PDCR: Perception-Decomposed Confidence Reward for Vision-Language ReasoningHee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom et al.CVPR 2026 · 3 citations
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity RecognitionShan Ning, Longtian Qiu, Jiaxuan Sun, Xuming HeCVPR 2026 · 1 citation
Builds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang et al.NeurIPS 2025 · 3 citations
- NoisyRollout: Reinforcing Visual Reasoning with Data AugmentationXiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du et al.NeurIPS 2025 · 104 citations
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 2 citations
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang et al.CVPR 2026 · 3 citations
- Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement LearningShuang Chen, Hangyu Guo, Zhaochen Su, Yafu Li et al.ICLR 2026 · 49 citations
