Lune

ICML2026Top-tier venue

Expanding the AI Evaluation Toolbox with Statistical Models

Drew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao, Julia Sharp, A. Bergman

2026Year
4Citations

Abstract

Benchmarks are widely used to evaluate and compare the performance of artificial intelligence systems. However, some approaches to computing benchmark metrics produce invalid uncertainty estimates or make unrecognized assumptions about the evaluation setting. We leverage statistical modeling to make two contributions to the practice of AI benchmarking. First, we formally distinguish measurements of benchmark accuracy (performance conditioned on a fixed benchmark) from generalized accuracy (performance on all potential test items similar to those included in the benchmark). Then, in a simulated setting and with large-scale evaluation of 22 API-access frontier large language models on 3 popular benchmarks, we show how analysis via generalized linear mixed models can accurately estimate generalized accuracy while more efficiently quantifying uncertainty compared to existing regression-free approaches. We also show how this approach can equip evaluators with important context on evaluation results, including variance decomposition and item difficulty estimates that illuminate important aspects of LLM performance and benchmark construction. Our findings highlight the benefits of explicitly specifying a statistical model for more accurate and reliable AI benchmarking.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c9133a57-7733-433d-a8af-d46068429c8e

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines