Lune

ACL2026顶会

Ranking Reasoning LLMs under Test-Time Scaling

Mohsen Hariri, Michael Hinczewski, Jing Ma, Vipin Chaudhary

2026年份
2被引次数
1顶会引用

摘要

Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking models in this regime remains underexplored. We formalize dense benchmark ranking under test-time scaling and introduce Scorio, a library that implements statistical ranking methods such as paired-comparison models, item response theory (IRT) models, voting rules, and graph- and spectral-based methods. Across 2020 reasoning models on four Olympiad-style math benchmarks (AIME'24, AIME'25, HMMT'25, and BrUMO'25; up to N=80N=80 trials), most full-trial rankings agree closely with the Bayesian gold standard BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 (mean Kendall's τb=0.93\tau_b = 0.93--0.950.95), and 1919--3434 methods recover exactly the same ordering. In the single-trial regime, the best methods reach τb≈0.86\tau_b \approx 0.86. Using greedy decoding as an empirical prior (BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N) reduces variance at N=1N=1 by 1616--52%52\%, but can bias rankings when greedy and stochastic sampling disagree. These results identify reliable ranking methods for both high- and low-budget test-time scaling. We release Scorio as an open-source library at https://github.com/mohsenhariri/scorio.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖