Lune

NeurIPS2025Top-tier venue

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation

Longtian Qiu, Shan Ning, Jiaxuan Sun, Xuming He

2025Year
6Citations
2Top-tier citations

Abstract

Reinforcement learning (RL) has shown promise in enhancing the general Chainof-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distribution. To address this, we propose NoisyGRPO, a systematic multimodal RL framework that introduces controllable noise into visual inputs for enhanced exploration and explicitly models the advantage estimation process via a Bayesian framework. Specifically, NoisyGRPO improves RL training by: (1) Noise-Injected Exploration Policy: Perturbing visual inputs with Gaussian noise to encourage exploration across a wider range of visual scenarios; and (2) Bayesian Advantage Estimation: Formulating advantage estimation as a principled Bayesian inference problem, where the injected noise level serves as a prior and the observed trajectory reward as the likelihood. This Bayesian modeling fuses both sources of information to compute a robust posterior estimate of trajectory advantage, effectively guiding MLLMs to prefer visually grounded trajectories over noisy ones. Experiments on standard CoT quality, general capability, and hallucination benchmarks demonstrate that NoisyGRPO substantially improves generalization and robustness, especially in RL settings with small-scale MLLMs such as Qwen2.5-VL 3B. The project page is available at https://artanic30.github.io/project_pages/NoisyGRPO.

2 Related Works Multi-modal Large Language Model Vision-language models (VLMs) have rapidly advanced in their ability to comprehend and reason over both visual and textual modalities and demonstrate great improvements over downstream tasks [35,64,12,9,31,40]. These models typically integrate visual encoders with large language models (LLMs) to enable cross-modal understanding and inference. Foundational works such as Flamingo [2] and BLIP-2 [20] established effective strategies for aligning vision and language components. The LLaVA-series [25,24,18] and SPHINXseries [22,10] further advanced the field by introducing visual instruction tuning, significantly enhancing multimodal capabilities. Large-scale models like GPT-4o [14] and Gemini [42] have demonstrated strong general-purpose visual understanding through large-scale multimodal pertaining. To address scalability, mixture-of-experts approaches-such as DeepSeek-VL2 [55], Uni-MoE [21], and MoVA [65]-improve computational efficiency by selectively activating expert modules based on input characteristics. More recently, domain-specific methods such as Math-LLaVA [47] and MAVIS [60] employed mathematical visual instruction tuning to improve VLMs' ability to interpret and solve complex multimodal math problems. At the same time, unified architectures like SEED-X [11], Chameleon [48], Show-O [56], and the Janus series [54, 29, 6] integrate both visual understanding and generation capabilities within a single framework. Despite these advances, most existing VLMs still struggle with robust visual reasoning, particularly in tasks that demand deep visual analysis and complex multi-step reasoning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 52bffe01-13b1-4bbf-830c-4d77352bd0cf

Cited by top-tier papers2

Ask how each one uses it

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines