Lune

ICML2026Top-tier venue

On the Power of (Approximate) Reward Models for Inference-Time Scaling: Sequential Monte Carlo and Beyond

Youheng Zhu, Yiping Lu

2026Year

Abstract

Inference-time scaling has recently emerged as a powerful paradigm for improving the reasoning capability of large language models. Among various approaches, Sequential Monte Carlo (SMC) has become a particularly important framework, enabling iterative generation, evaluation, rejection, and resampling of intermediate reasoning trajectories. A central component in this process is the reward model, which evaluates partial solutions and guides the allocation of computation during inference. However, in practice, true reward models are never available. All deployed systems rely on approximate reward models, raising a fundamental question: Why and when do approximate reward models suffice for effective inference-time scaling? In this work, we provide a theoretical answer. We identify the Bellman error of the approximate reward model as the key quantity governing the effectiveness of SMC-based inference-time scaling. For a reasoning process of length TT, we show that if the Bellman error of the approximate reward model is bounded by O(1/T)O(1/T), then combining this reward model with SMC reduces the computational complexity of reasoning from exponential in TT to polynomial in TT. This yields an exponential improvement in inference efficiency despite using only approximate rewards.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines