: A Generalist Value Model for Any Policy at State Zero
Yi-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun, Qi GU, Hui Su, Xunliang Cai, De-Chuan Zhan, Han-Jia Ye
摘要
Policy gradient methods rely on a baseline to measure the relative advantage of an action, ensuring the model reinforces behaviors that outperform its current average capability. In the training of Large Language Models (LLMs) using Actor-Critic methods (e.g., PPO), this baseline is typically estimated by a Value Model (Critic) often as large as the policy model itself. However, as the policy continuously evolves, the value model requires expensive, synchronous incremental training to accurately track the shifting capabilities of the policy. To avoid this overhead, Group Relative Policy Optimization (GRPO) eliminates the coupled value model by using the average reward of a group of rollouts as the baseline; yet, this approach necessitates extensive sampling to maintain estimation stability. In this paper, we propose V 0 , a Generalist Value Model capable of estimating the expected performance of any model on unseen prompts without requiring parameter updates. We reframe value estimation by treating the policy's dynamic capability as an explicit context input; specifically, we leverage a history of instruction-performance pairs to dynamically profile the model, departing from the traditional paradigm that relies on parameter fitting to perceive capability shifts. Focusing on value estimation at State Zero (i.e., the initial prompt, hence V 0 ), our model serves as a critical resource scheduler. During GRPO training, V 0 predicts success rates prior to rollout, allowing for efficient sampling budget allocation; during deployment, it functions as a router, dispatching instructions to the most cost-effective and suitable model. Empirical results demonstrate that V 0 significantly outperforms heuristic budget allocation and achieves a Pareto-optimal trade-off between performance and cost in LLM routing tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- AlphaMath Almost Zero: Process Supervision without ProcessGuoxin Chen, Minpeng Liao, Chengxi Li, Kai FanNeurIPS 2024 · 被引用 219 次
相关 Paper
- End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement LearningGuanzhong Chen, Shaoxiong Yang, Chao Li, Wei Liu 等ACL 2026 · 被引用 8 次
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksTianyi Wang, Yixia Li, Long Li, Yibiao Chen 等ACL 2026 · 被引用 8 次
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic TasksShuo He, Lang Feng, Qi Wei, Xin Cheng 等ICLR 2026 · 被引用 36 次
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- Reinforcement Learning for Large Language Models via Group Preference Reward ShapingHuaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao 等EMNLP 2025
