On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
Yue Yu, Qiwei Di, Quanquan Gu, Dongruo Zhou
Abstract
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of-n (BoN) sampling and sequential revision, their fundamental limits remain unclear. We address this gap by analyzing a mixture-of-reference policy model and proving that standard BoN is inherently suboptimal. To move closer to the optimal frontier, we study reward-filtered sequential inference, a simple procedure that selectively incorporates only high-reward generations into the context. This mechanism concentrates computation on superior policy candidates and suppresses inferior ones. On the theoretical side, we show that reward-filtered sequential inference yields strictly stronger guarantees than standard TTC paradigms. On the empirical side, we evaluate such an inference strategy across diverse benchmarks and observe consistent improvements over widely used approaches, demonstrating the practical effectiveness of our framework. How to effectively utilize large language models (LLMs) for solving new tasks has become a central research question. Among the many approaches, Test-Time Compute (TTC) has recently attracted significant attention. The key idea of TTC is to allocate additional computation during inference to improve task performance. Unlike post-training approaches such as fine-tuning or reinforcement learning, TTC requires no additional training of the base model. As a result, inference-time alignment methods provide a lightweight yet powerful alternative that greatly simplifies deployment. Wellknown TTC methods include Best-of-N (BoN) sampling, chain-of-thought (CoT) reasoning, and their many variants (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84abeee6-2903-477b-8948-ba781918dcfaBuilds on32
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Diversity Matters: Revisiting Test-Time Compute in Vision-Language ModelsYijie Tong, Yifan Hou, Shaobo Cui, Antoine Bosselut et al.ICML 2026 · 1 citation
- Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and DistillationJuno Kim, Denny Wu, Jason D. Lee, Taiji SuzukiICML 2025
- Test-time Prompt InterventionChenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao et al.AAAI 2026 · 8 citations
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time AlignmentAudrey Huang, Adam Block, Qinghua Liu, Nan Jiang et al.ICML 2025
- Let's (not) just put things in Context: Test-time Training for Long-context LLMsRachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan et al.ICLR 2026 · 20 citations
