metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Alexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric Schulz
Abstract
Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either correctly or incorrectly). However, high correlations within and between benchmark scores suggest that (1) there exists a small set of common underlying abilities that these benchmarks measure, and ( 2 ) items tap into redundant information and the benchmarks may thus be considerably compressed. We use data from n > 5000 LLMs to identify the most informative items of six benchmarks, ARC, GSM8K, HellaSwag, MMLU, TruthfulQA and WinoGrande (with d = 28, 632 items in total). From them we distill a sparse benchmark, metabench, that has less than 3% of the original size of all six benchmarks combined. This new sparse benchmark goes beyond point scores by yielding estimators of the underlying benchmark-specific abilities. We show that these estimators (1) can be used to reconstruct each original individual benchmark score with, on average, 1.24% root mean square error (RMSE), (2) reconstruct the original total score with 0.58% RMSE, and (3) have a single underlying common factor whose Spearman correlation with the total score is r = 0.94.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e915f48-6599-4a0c-a0a2-bf27ff46df6dCited by top-tier papers14
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng et al.NeurIPS 2025 · 160 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static BenchmarksPeiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng et al.ICML 2026 · 23 citations
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response TheoryHongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han et al.AAAI 2026 · 15 citations
- Predicting LLM Reasoning Performance with Small Proxy ModelWoosung Koh, Juyoung Suk, Sungjun Han, Se-Young Yun et al.ICLR 2026 · 5 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM EvaluationZhuo Wang, Wen Wu, Guoqing Wang, Guangze Ye et al.AAAI 2026 · 1 citation
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- SparseEval: Efficient Evaluation of Large Language Models by Sparse OptimizationTaolin Zhang, Hang Guo, Wang Lu, Tao Dai et al.ICLR 2026 · 6 citations
- Learning More from Less: Unlocking Internal Representations for Benchmark CompressionYueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan et al.ICML 2026
