Lune

NeurIPS2025顶会

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation

Longtian Qiu, Shan Ning, Jiaxuan Sun, Xuming He

2025年份
6被引次数
2顶会引用

摘要

Reinforcement learning (RL) has shown promise in enhancing the general Chainof-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distribution. To address this, we propose NoisyGRPO, a systematic multimodal RL framework that introduces controllable noise into visual inputs for enhanced exploration and explicitly models the advantage estimation process via a Bayesian framework. Specifically, NoisyGRPO improves RL training by: (1) Noise-Injected Exploration Policy: Perturbing visual inputs with Gaussian noise to encourage exploration across a wider range of visual scenarios; and (2) Bayesian Advantage Estimation: Formulating advantage estimation as a principled Bayesian inference problem, where the injected noise level serves as a prior and the observed trajectory reward as the likelihood. This Bayesian modeling fuses both sources of information to compute a robust posterior estimate of trajectory advantage, effectively guiding MLLMs to prefer visually grounded trajectories over noisy ones. Experiments on standard CoT quality, general capability, and hallucination benchmarks demonstrate that NoisyGRPO substantially improves generalization and robustness, especially in RL settings with small-scale MLLMs such as Qwen2.5-VL 3B. The project page is available at https://artanic30.github.io/project_pages/NoisyGRPO.

2 Related Works Multi-modal Large Language Model Vision-language models (VLMs) have rapidly advanced in their ability to comprehend and reason over both visual and textual modalities and demonstrate great improvements over downstream tasks [35,64,12,9,31,40]. These models typically integrate visual encoders with large language models (LLMs) to enable cross-modal understanding and inference. Foundational works such as Flamingo [2] and BLIP-2 [20] established effective strategies for aligning vision and language components. The LLaVA-series [25,24,18] and SPHINXseries [22,10] further advanced the field by introducing visual instruction tuning, significantly enhancing multimodal capabilities. Large-scale models like GPT-4o [14] and Gemini [42] have demonstrated strong general-purpose visual understanding through large-scale multimodal pertaining. To address scalability, mixture-of-experts approaches-such as DeepSeek-VL2 [55], Uni-MoE [21], and MoVA [65]-improve computational efficiency by selectively activating expert modules based on input characteristics. More recently, domain-specific methods such as Math-LLaVA [47] and MAVIS [60] employed mathematical visual instruction tuning to improve VLMs' ability to interpret and solve complex multimodal math problems. At the same time, unified architectures like SEED-X [11], Chameleon [48], Show-O [56], and the Janus series [54, 29, 6] integrate both visual understanding and generation capabilities within a single framework. Despite these advances, most existing VLMs still struggle with robust visual reasoning, particularly in tasks that demand deep visual analysis and complex multi-step reasoning.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 52bffe01-13b1-4bbf-830c-4d77352bd0cf

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖