Don’t Pass@k: A Bayesian Framework for Large Language Model Evaluation
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin Chaudhary
摘要
Pass@ is widely used to report performance for LLM reasoning, but it often yields unstable, misleading rankings, especially when the number of trials (samples) is limited and compute is constrained. We present a principled Bayesian evaluation framework that replaces Pass@ and average accuracy over trials (avg@) with posterior estimates of a model's underlying success probability and credible intervals, yielding stable rankings and a transparent decision rule for differences. Evaluation outcomes are modeled as categorical (not just 0/1) with a Dirichlet prior, giving closed-form expressions for the posterior mean and uncertainty of any weighted rubric and enabling the use of prior evidence when appropriate. Theoretically, under a uniform prior, the Bayesian posterior mean is order-equivalent to average accuracy (Pass@), explaining its empirical robustness while adding principled uncertainty. Empirically, in simulations with known ground-truth success rates and on AIME'24/'25, HMMT, and BrUMO, the Bayesian/avg procedure achieves faster convergence and greater rank stability than Pass@ and recent variants, enabling reliable comparisons at far smaller sample counts. The framework clarifies when observed gaps are statistically meaningful (non-overlapping credible intervals) versus noise, and it naturally extends to graded, rubric-based evaluations. Together, these results recommend replacing Pass@ for LLM evaluation and ranking with a posterior-based, compute-efficient protocol that unifies binary and non-binary evaluation while making uncertainty explicit. Source code is available at https://github.com/mohsenhariri/scorio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ranking Reasoning LLMs under Test-Time ScalingMohsen Hariri, Michael Hinczewski, Jing Ma, Vipin ChaudharyACL 2026 · 被引用 2 次
- Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative SignalsZihan Dong, Zhixian Zhang, Yang Zhou, Can Jin 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World DataAlana Renda, Jillian Ross, Jacob AndreasICLR 2026 · 被引用 3 次
- Efficient Bayesian Inference from Noisy Pairwise ComparisonsTill Aczel, Lucas Theis, Roger WattenhoferICML 2026 · 被引用 2 次
- Optimal Bayesian Stopping for Efficient Inference of Consistent LLM AnswersJingkai Huang, Will Ma, Zhengyuan ZhouICML 2026 · 被引用 3 次
- Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian AlternativeKoen OostermeijerICML 2026
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
