Lune

NeurIPS2025Top-tier venue

Language Models can Self-Improve at State-Value Estimation for Better Search

Ethan Mendes, Alan Ritter

2025Year
5Citations
1Top-tier citations

Abstract

Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, particularly in interactive domains such as web tasks. We introduce Self-Taught Lookahead (STL), a reward-free framework that improves language model-based value functions by reasoning explicitly about state transitions. STL can be viewed as a chain-of-thought analogue of the value iteration algorithm: instead of regressing directly on numeric values, a value LLM is trained to simulate a step of lookahead in natural language-predicting the next action, resulting state, and rationale for its value, thereby refining value estimates without any labeled data. This self-supervised procedure yields more accurate statevalue predictions, which in turn enable lightweight search algorithms to expand fewer states while maintaining strong performance. Empirically, STL-trained value models built on moderately sized (8B parameter) open-weight LLMs boost web agent success rates by 39%, achieving comparable performance with proprietary models. STL also generalizes to multi-hop QA and math puzzles. We find that STL enables small open-source models to guide efficient search, reducing inference costs by integrating explicit reasoning with value learning. In this paper, we introduce Self-Taught Lookahead (STL), a Reward and Demo Free , selfsupervised framework for interactive, multi-step reasoning tasks. Building on evidence that search quality is strongly influenced by the accuracy of the state-value estimator [10, 37] , STL improves an 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines