Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning
Zizhe Chen, Jiqian Dong, Yizhou Tian, Garry YANG, Yongqiang Chen, Zhitang Chen, James Cheng
Abstract
Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state value estimation is critical for stable training in classical RL, it remains an underexplored challenge in LLM post-training. In this work, we introduce the State Value Estimation Benchmark (SVEB) to assess state estimation within existing RL frameworks and show that critics in standard approaches like PPO collapse to a coarse group-average baseline. To address this, we propose two techniques: Numca, which leverages numerical spans as gradable milestones for state value estimation, and Hista, a framework that uses LLM's hidden states as representation to weighted average disjoint rollouts and their return. Extensive experiments demonstrate that both methods yield more accurate state value estimates and enhance training performance across different RL algorithms and model sizes without incurring significant computational overhead. Code available at https://github. com/VOXXXX1874/Hista . How does accurate state value improve LLM RL training, and how can it be estimated effectively? To answer this question, we first construct a State Value Estimation Benchmark (SVEB) to quantify the discrepancy
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f69b3ec2-d87f-433b-998e-e3af63282de6Builds on19
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
Related papers
- Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable RewardsHaechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang et al.ICML 2026
- : A Generalist Value Model for Any Policy at State ZeroYi-Kai Zhang, Zhiyuan Yao, Hongyan Hao, Yueqing Sun et al.ICML 2026 · 3 citations
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu et al.ICLR 2026 · 5 citations
- Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language ModelsAshutosh Baheti, Ximing Lu, Faeze Brahman, Ronan Le Bras et al.ICLR 2024 · 16 citations
- From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent TrainingZishang Jiang, tingyun li, Jinyi Han, Xinyi Wang et al.ICML 2026
