Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning
Zizhe Chen, Jiqian Dong, Yizhou Tian, Garry YANG, Yongqiang Chen, Zhitang Chen, James Cheng
摘要
Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state value estimation is critical for stable training in classical RL, it remains an underexplored challenge in LLM post-training. In this work, we introduce the State Value Estimation Benchmark (SVEB) to assess state estimation within existing RL frameworks and show that critics in standard approaches like PPO collapse to a coarse group-average baseline. To address this, we propose two techniques: Numca, which leverages numerical spans as gradable milestones for state value estimation, and Hista, a framework that uses LLM's hidden states as representation to weighted average disjoint rollouts and their return. Extensive experiments demonstrate that both methods yield more accurate state value estimates and enhance training performance across different RL algorithms and model sizes without incurring significant computational overhead. Code available at https://github. com/VOXXXX1874/Hista . How does accurate state value improve LLM RL training, and how can it be estimated effectively? To answer this question, we first construct a State Value Estimation Benchmark (SVEB) to quantify the discrepancy
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable RewardsHaechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang 等ICML 2026
- : A Generalist Value Model for Any Policy at State ZeroYi-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun 等ICML 2026 · 被引用 3 次
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu 等ICLR 2026 · 被引用 5 次
- Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language ModelsAshutosh Baheti, Ximing Lu, Faeze Brahman, Ronan Le Bras 等ICLR 2024 · 被引用 16 次
- From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent TrainingZishang Jiang, tingyun li, Jinyi Han, Xinyi Wang 等ICML 2026
