Active Evaluation Acquisition for Efficient LLM Benchmarking
Yang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba, Graham Horwood
Abstract
As large language models (LLMs) become increasingly versatile, numerous large scale benchmarks have been developed to thoroughly assess their capabilities. These benchmarks typically consist of diverse datasets and prompts to evaluate different aspects of LLM performance. However, comprehensive evaluations on hundreds or thousands of prompts incur tremendous costs in terms of computation, money, and time. In this work, we investigate strategies to improve evaluation efficiency by selecting a subset of examples from each benchmark using a learned policy. Our approach models the dependencies across test examples, allowing accurate prediction of the evaluation outcomes for the remaining examples based on the outcomes of the selected ones. Consequently, we only need to acquire the actual evaluation outcomes for the selected subset. We rigorously explore various subset selection policies and introduce a novel RL-based policy that leverages the captured dependencies. Empirical results demonstrate that our approach significantly reduces the number of evaluation prompts required while maintaining accurate performance estimates compared to previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2a50ea6-b2b0-4082-885f-c59e09d167ddCited by top-tier papers9
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- DISCO: Diversifying Sample Condensation for Efficient Model EvaluationAlexander Rubinstein, Benjamin Raible, Martin Gubri, Seong Joon OhICLR 2026 · 4 citations
- Metritocracy: Representative Metrics for Lite BenchmarksAriel D. Procaccia, Ben Schiffer, Serena Wang, Shirley ZhangNeurIPS 2025 · 2 citations
- MetaEval: Measuring the Discrimination of Benchmarks for Efficient LLM EvaluationZhuo Wang, Wen Wu, Guoqing Wang, Guangze Ye et al.AAAI 2026 · 1 citation
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 163 citations
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 81 citations
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
Related papers
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM EvaluationGayathri Saranathan, Cong Xu, Mahammad Parwez Alam, Tarun Kumar et al.ACL 2025 · 3 citations
- Examining the robustness of LLM evaluation to the distributional assumptions of benchmarksCharlotte Siska, Katerina Marazopoulou, Melissa Ailem, James BonoACL 2024
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang et al.EMNLP 2024 · 6 citations
- LLM-Powered Benchmark Factory: Reliable, Generic, and EfficientPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ACL 2026 · 8 citations
