VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang
Abstract
Reinforcement Learning (RL) in real-world environments often suffers from ambiguous or incomplete reward supervision, which undermines policy stability and generalization. Such noise may cause models to ignore key information or even collapse in advantage estimation. We find that a strong value model is essential for absorbing unstable signals and producing reliable advantages, offering denser and more robust supervision than the reward model. To better optimize noisy supervision, we propose VRPO, a framework that enhances value modeling for robust RL in LLM post-training. VRPO integrates (1) auxiliary losses guided by entropy and perplexity from a frozen language model, and (2) a variational information bottleneck, enabling the value model to filter noise and capture key words. This design allows the value model to correct noise rewards and generate more reliable advantage estimates, transforming it from a passive predictor into an active noise regulator. Experiments on multi-turn dialogue, math reasoning, and science QA with both rule-based and model-based rewards show that VRPO consistently outperforms baselines such as PPO and GRPO. Our work highlight the central role of the value model in Robust RL and provide a principled and practical approach to policy optimization under noisy supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b482eb7-80f7-4006-8d1b-04490e8f74dfBuilds on9
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingYuchun Miao, Sen Zhang, Liang Ding, Rong Bao et al.NeurIPS 2024 · 108 citations
- RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesJie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao et al.ICML 2024 · 42 citations
- MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou et al.ACL 2022 · 35 citations
Related papers
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang et al.CVPR 2026 · 4 citations
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian et al.ICLR 2026 · 8 citations
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian EstimationLongtian Qiu, Shan Ning, Jiaxuan Sun, Xuming HeNeurIPS 2025 · 6 citations
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoningJiashun Liu, Johan S. Obando-Ceron, Han Lu, Yancheng He et al.ICLR 2026 · 11 citations
- VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision–Language ModelsXUEGE HOU, Wenshuo Li, Yali Li, Han Shu et al.CVPR 2026
