VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang
摘要
Reinforcement Learning (RL) in real-world environments often suffers from ambiguous or incomplete reward supervision, which undermines policy stability and generalization. Such noise may cause models to ignore key information or even collapse in advantage estimation. We find that a strong value model is essential for absorbing unstable signals and producing reliable advantages, offering denser and more robust supervision than the reward model. To better optimize noisy supervision, we propose VRPO, a framework that enhances value modeling for robust RL in LLM post-training. VRPO integrates (1) auxiliary losses guided by entropy and perplexity from a frozen language model, and (2) a variational information bottleneck, enabling the value model to filter noise and capture key words. This design allows the value model to correct noise rewards and generate more reliable advantage estimates, transforming it from a passive predictor into an active noise regulator. Experiments on multi-turn dialogue, math reasoning, and science QA with both rule-based and model-based rewards show that VRPO consistently outperforms baselines such as PPO and GRPO. Our work highlight the central role of the value model in Robust RL and provide a principled and practical approach to policy optimization under noisy supervision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingYuchun Miao, Sen Zhang, Liang Ding, Rong Bao 等NeurIPS 2024 · 被引用 108 次
- RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesJie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao 等ICML 2024 · 被引用 42 次
- MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou 等ACL 2022 · 被引用 35 次
相关 Paper
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM TrainingDonglai Xu, Hongzheng Yang, Yuzhi Zhao, Pingping Zhang 等CVPR 2026 · 被引用 4 次
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian 等ICLR 2026 · 被引用 8 次
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian EstimationLongtian Qiu, Shan Ning, Jiaxuan Sun, Xuming HeNeurIPS 2025 · 被引用 6 次
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoningJiashun Liu, Johan S. Obando-Ceron, Han Lu, Yancheng He 等ICLR 2026 · 被引用 11 次
- VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision–Language ModelsXUEGE HOU, Wenshuo Li, Yali Li, Han Shu 等CVPR 2026
