ROC-n-reroll: How verifier imperfection affects test-time scaling
Florian E. Dorner, Yatong Chen, André F Cruz, Fanny Yang
Abstract
Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (BoN) and Rejection Sampling (RS) that make use of a verifier to enable test-time scaling. However, to date there is little theoretical understanding of how verifier imperfection affects performance -a gap we address in this work. Specifically, we prove that the instance-level accuracy of these methods is precisely characterized by the geometry of the verifier's ROC curve. Our theory has two important takeaways, confirmed by experiments with Qwen and LLama models on GSM8K and MATH500. First, RS outperforms BoN for fixed compute, while both methods converge to the same accuracy in the infinite-compute limit. Second, it is generally impossible to predict the high-compute performance of either method based on observations in the low-compute regime. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7262f9bf-b64d-4c20-95c3-241d2a1c8e18Cited by top-tier papers2
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimalityArpan Mukherjee, Marcello Bullo, Debabrota Basu, Deniz GunduzICLR 2026 · 3 citations
- Large Language Models Explore by Latent DistillingYuanhao Zeng, Ao Lu, Lufei Li, Zheng Zhang et al.ICML 2026
Builds on18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingYiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang et al.NeurIPS 2025 · 66 citations
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding et al.NeurIPS 2025 · 16 citations
- Sample Complexity and Representation Ability of Test-time Scaling ParadigmsBaihe Huang, Shanda Li, Tianhao Wu, Yiming Yang et al.ICLR 2026 · 11 citations
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for ReasoningCharlie Victor Snell, Jaehoon Lee, Kelvin Xu, Aviral KumarICLR 2025
- T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language ModelsMinki Kang, Jongwon Jeong, Jaewoong ChoICLR 2026 · 14 citations
