On the Power of (Approximate) Reward Models for Inference-Time Scaling: Sequential Monte Carlo and Beyond
Youheng Zhu, Yiping Lu
摘要
Inference-time scaling has recently emerged as a powerful paradigm for improving the reasoning capability of large language models. Among various approaches, Sequential Monte Carlo (SMC) has become a particularly important framework, enabling iterative generation, evaluation, rejection, and resampling of intermediate reasoning trajectories. A central component in this process is the reward model, which evaluates partial solutions and guides the allocation of computation during inference. However, in practice, true reward models are never available. All deployed systems rely on approximate reward models, raising a fundamental question: Why and when do approximate reward models suffice for effective inference-time scaling? In this work, we provide a theoretical answer. We identify the Bellman error of the approximate reward model as the key quantity governing the effectiveness of SMC-based inference-time scaling. For a reasoning process of length , we show that if the Bellman error of the approximate reward model is bounded by , then combining this reward model with SMC reduces the computational complexity of reasoning from exponential in to polynomial in . This yields an exponential improvement in inference efficiency despite using only approximate rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based DecodingXiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia 等NeurIPS 2025 · 被引用 147 次
相关 Paper
- Know What You Don't Know: Uncertainty Calibration of Process Reward ModelsYoung-Jin Park, Kristjan Greenewald, Kaveh Alimohammadi, Hao Wang 等NeurIPS 2025 · 被引用 21 次
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time ScalingIndranil Halder, Cengiz PehlevanICML 2026 · 被引用 4 次
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar 等NeurIPS 2025 · 被引用 8 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward ModelingJiachun Li, Pengfei Cao, Zhuoran Jin, Yubo Chen 等ICLR 2026 · 被引用 3 次
