Expanding the AI Evaluation Toolbox with Statistical Models
Drew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao, Julia Sharp, A. Bergman
Abstract
Benchmarks are widely used to evaluate and compare the performance of artificial intelligence systems. However, some approaches to computing benchmark metrics produce invalid uncertainty estimates or make unrecognized assumptions about the evaluation setting. We leverage statistical modeling to make two contributions to the practice of AI benchmarking. First, we formally distinguish measurements of benchmark accuracy (performance conditioned on a fixed benchmark) from generalized accuracy (performance on all potential test items similar to those included in the benchmark). Then, in a simulated setting and with large-scale evaluation of 22 API-access frontier large language models on 3 popular benchmarks, we show how analysis via generalized linear mixed models can accurately estimate generalized accuracy while more efficiently quantifying uncertainty compared to existing regression-free approaches. We also show how this approach can equip evaluators with important context on evaluation results, including variance decomposition and item difficulty estimates that illuminate important aspects of LLM performance and benchmark construction. Our findings highlight the benefits of explicitly specifying a statistical model for more accurate and reliable AI benchmarking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9133a57-7733-433d-a8af-d46068429c8eBuilds on6
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani et al.ACL 2025 · 144 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
Related papers
- AI Cartography: Mapping the Latent Landscape of AI Benchmark EcosystemsMichael Hardy, Anka Reuel, Lijin Zhang, Jodi Casabianca et al.ICML 2026
- On the Evaluation of Capability Estimation Methods for Large Language ModelsQiang Hu, Jin Wen, Yao Zhang, Maxime Cordy et al.AAAI 2026
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 40 citations
- Benchmarking at the Edge of ComprehensionSamuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb et al.ICML 2026
