Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Omar Elmansouri, Fathinah Izzati, Mohamed El Amine Seddik, Salem Lahlou
Abstract
Reinforcement learning from human feedback (RLHF) or verifiable rewards (RLVR), the standard paradigm for aligning LLMs or building recent SOTA reasoning models, is highly sensitive to noise from inconsistent or erroneous rewards. Yet, the interaction between such noise and widely used group-based policy optimization methods remains underexplored. We introduce a noise-robust Group Relative Policy Optimization (GRPO) and Done Right GRPO (Dr.GRPO) framework that explicitly models reward corruption as Bernoulli noise. Our method applies noise correction after estimating reward flip probabilities to debias the learning signal, yielding unbiased gradient estimates. Theoretical analysis shows that group-based methods inherently mitigate individual-level noise, and our correction strategy amplifies this robustness. Empirically, we observe consistent improvements across math and code tasks when applying our noise correction to standard reward model usage, with particular gains of up to 6.7 percentage points in accuracy on math tasks and 1.5 on code tasks under realistic reward model conditions. This work bridges label-noise correction from supervised learning with modern RLHF, offering both theoretical insights and a practical algorithm for noisy real-world deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy TrainingYoussef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure et al.ICLR 2026 · 39 citations
- Robust Reinforcement Learning from Corrupted Human FeedbackAlexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li et al.NeurIPS 2024 · 30 citations
Related papers
- AG-GRPO: Answer-Guided GRPO for Masked Diffusion Language ModelsJuhyeong Kim, Gyunyeop Kim, Sangwoo KangACL 2026
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage ShapingThanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Dac Lai et al.ICLR 2026 · 55 citations
- Spurious Rewards: Rethinking Training Signals in RLVRRulin Shao, Stella Li, Rui Xin, Scott Geng et al.ICML 2026
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang et al.CVPR 2026 · 4 citations
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine et al.ICLR 2026 · 218 citations
