Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
Arpan Mukherjee, Marcello Bullo, Debabrota Basu, Deniz Gunduz
Abstract
While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), role of the verifier and its imperfections remain underexplored. The effect of verification manifests through interactions of three quantities: (i) the generator’s coverage, (ii) the verifier’s region of convergence (ROC), and (iii) the sampling algorithm’s sub-optimality. Though recent studies capture subsets of these factors, a unified framework quantifying the geometry of their interplay is missing. We frame verifiable test-time scaling as a transport problem. This characterizes the interaction of coverage, ROC, and sub-optimality, and uncovers that the sub-optimality-coverage curve exhibits three regimes. A transport regime — where sub-optimality increases with coverage, a policy improvement regime — where sub-optimality may decrease with coverage, depending on the verifier’s ROC, and a saturation regime — where sub-optimality plateaus, unaffected by coverage. We further propose and analyze two classes of sampling algorithms — sequential and batched, and examine how their computational complexities shape these trade-offs. Empirical results with Qwen, Llama, and Gemma models corroborate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdcee7b6-aa9c-4eb4-90b2-7a321ef36f7cCited by top-tier papers1
Ask how each one uses itBuilds on6
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingYiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang et al.NeurIPS 2025 · 66 citations
- ROC-n-reroll: How verifier imperfection affects test-time scalingFlorian E. Dorner, Yatong Chen, André F Cruz, Fanny YangICLR 2026 · 13 citations
- Sample Complexity and Representation Ability of Test-time Scaling ParadigmsBaihe Huang, Shanda Li, Tianhao Wu, Yiming Yang et al.ICLR 2026 · 11 citations
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time AlignmentAudrey Huang, Adam Block, Qinghua Liu, Nan Jiang et al.ICML 2025
Related papers
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsYefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh et al.ICLR 2026 · 17 citations
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsYuda Song, Hanlin Zhang, Carson Eisenach, Sham M. Kakade et al.ICLR 2025 · 3 citations
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding et al.NeurIPS 2025 · 16 citations
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
- Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time ScalingHao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo et al.NeurIPS 2025 · 8 citations
