Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
Arpan Mukherjee, Marcello Bullo, Debabrota Basu, Deniz Gunduz
摘要
While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), role of the verifier and its imperfections remain underexplored. The effect of verification manifests through interactions of three quantities: (i) the generator’s coverage, (ii) the verifier’s region of convergence (ROC), and (iii) the sampling algorithm’s sub-optimality. Though recent studies capture subsets of these factors, a unified framework quantifying the geometry of their interplay is missing. We frame verifiable test-time scaling as a transport problem. This characterizes the interaction of coverage, ROC, and sub-optimality, and uncovers that the sub-optimality-coverage curve exhibits three regimes. A transport regime — where sub-optimality increases with coverage, a policy improvement regime — where sub-optimality may decrease with coverage, depending on the verifier’s ROC, and a saturation regime — where sub-optimality plateaus, unaffected by coverage. We further propose and analyze two classes of sampling algorithms — sequential and batched, and examine how their computational complexities shape these trade-offs. Empirical results with Qwen, Llama, and Gemma models corroborate our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 被引用 273 次
- Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingYiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang 等NeurIPS 2025 · 被引用 66 次
- ROC-n-reroll: How verifier imperfection affects test-time scalingFlorian E. Dorner, Yatong Chen, André F Cruz, Fanny YangICLR 2026 · 被引用 13 次
- Sample Complexity and Representation Ability of Test-time Scaling ParadigmsBaihe Huang, Shanda Li, Tianhao Wu, Yiming Yang 等ICLR 2026 · 被引用 11 次
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time AlignmentAudrey Huang, Adam Block, Qinghua Liu, Nan Jiang 等ICML 2025
相关 Paper
- Variation in Verification: Understanding Verification Dynamics in Large Language ModelsYefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh 等ICLR 2026 · 被引用 17 次
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsYuda Song, Hanlin Zhang, Carson Eisenach, Sham M. Kakade 等ICLR 2025 · 被引用 3 次
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding 等NeurIPS 2025 · 被引用 16 次
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
- Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time ScalingHao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo 等NeurIPS 2025 · 被引用 8 次
