The Limits of Inference Scaling Through Resampling
Benedikt Stroebl, Sayash Kapoor, Arvind Narayanan
Abstract
Recent research has generated hope that inference scaling, such as resampling solutions until they pass verifiers like unit tests, could allow weaker models to match stronger ones. Beyond inference, this approach also enables training reasoning models, where data is curated using rejection sampling against a verifier. However, we show that this approach is fundamentally limited when verifiers are imperfect and have a non-zero probability of producing false positives. Resampling cannot decrease this probability, so it imposes an upper bound to the accuracy of resampling-based inference scaling, regardless of compute budget. Our analysis shows that there is a strong correlation between the model’s single-sample accuracy and its false positive rate on HumanEval and MBPP, whose unit tests have limited coverage. Therefore, no amount of inference scaling of weaker models can enable them to match the single-sample accuracy of a sufficiently strong model. Empirical results show that optimal sampling attempts are often fewer than 10, as the negative utility of false positives outweighs benefits, bending inference scaling curves downward. Finally, false positives may have other undesirable qualities, like poor adherence to coding style conventions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9b1c6b0-e66e-4038-b7a5-3fac3afc85a2Cited by top-tier papers6
- Best-of-N JailbreakingJohn Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer et al.NeurIPS 2025 · 78 citations
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang et al.NeurIPS 2025 · 33 citations
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng et al.NeurIPS 2025 · 20 citations
- Bootstrapping Hierarchical Autoregressive Formal Reasoner with Chain-of-Proxy-AutoformalizationQi Liu, Xinhao Zheng, Renqiu Xia, Qinxiang Cao et al.NeurIPS 2025 · 3 citations
- Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-SolvingQi Liu, Xinhao Zheng, Renqiu Xia, Xingzhi Qi et al.ICML 2026
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling VerificationEric Zhao, Pranjal Awasthi, Sreenivas GollapudiICML 2025
- ROC-n-reroll: How verifier imperfection affects test-time scalingFlorian E. Dorner, Yatong Chen, André F Cruz, Fanny YangICLR 2026 · 13 citations
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 87 citations
- Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy VerifierYegor Denisov-Blanch, Joshua Kazdan, Jessica Chudnovsky, Rylan Schaeffer et al.ICML 2026
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu et al.NeurIPS 2025 · 43 citations
