tinyBenchmarks: evaluating LLMs with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, Mikhail Yurochkin
Abstract
The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities. These benchmarks consist of tens of thousands of examples making evaluation of LLMs very expensive. In this paper, we investigate strategies to reduce the number of evaluations needed to assess the performance of an LLM on several key benchmarks. For example, we show that to accurately estimate the performance of an LLM on MMLU, a popular multiple-choice QA benchmark consisting of 14K examples, it is sufficient to evaluate this LLM on 100 curated examples. We release evaluation tools and tiny versions of popular benchmarks: Open LLM Leaderboard, MMLU, HELM, and AlpacaEval 2.0. Our empirical analysis demonstrates that these tools and tiny benchmarks are sufficient to reliably and efficiently reproduce the original evaluation results 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0ea98b1-f94d-44ce-9d77-a3323e010c7eCited by top-tier papers68
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen et al.ICLR 2026 · 69 citations
- On the Worst Prompt Performance of Large Language ModelsBowen Cao, Deng Cai, Zhisong Zhang, Yuexian Zou et al.NeurIPS 2024 · 65 citations
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 53 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
Related papers
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba et al.ICML 2025
- metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language ModelsAlexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric SchulzICLR 2025
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 7 citations
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
- LLM-Powered Benchmark Factory: Reliable, Generic, and EfficientPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ACL 2026 · 8 citations
