How Benchmark Prediction from Fewer Data Misses the Mark
Guanhua Zhang, Florian E. Dorner, Moritz Hardt
摘要
Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets. Benchmark prediction (also called efficient LLM evaluation) aims to select a small subset of evaluation points and predict overall benchmark performance from that subset. In this paper, we systematically assess the strengths and limitations of 11 benchmark prediction methods across 19 diverse benchmarks. First, we identify a highly competitive baseline: Take a random sample and fit a regression model on the sample to predict missing entries. Outperforming most existing methods, this baseline challenges the assumption that careful subset selection is necessary for benchmark prediction. Second, we discover that all existing methods crucially depend on model similarity. They work best when interpolating scores among similar models. The effectiveness of benchmark prediction sharply declines when new models have higher accuracy than previously seen models. In this setting of extrapolation, none of the previous methods consistently beat a simple average over random samples. To improve over the sample average, we introduce a new method inspired by augmented inverse propensity weighting. This method consistently outperforms the random sample average even for extrapolation. However, its performance still relies on model similarity and the gains are modest in general. This shows that benchmark prediction fails just when it is most needed: at the evaluation frontier, where the goal is to evaluate new models of unknown capabilities † .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 被引用 7 次
- DISCO: Diversifying Sample Condensation for Efficient Model EvaluationAlexander Rubinstein, Benjamin Raible, Martin Gubri, Seong Joon OhICLR 2026 · 被引用 4 次
- FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random SequencesRichardeau Gurvan, Gohar Dashyan, Erwan Le Merrer, Gilles TredanICML 2026 · 被引用 2 次
- Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm That Provably Exploits Model SimilarityZifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang 等ICML 2026
- Putting HUMANS first: Efficient LAM Evaluation with Human Preference AlignmentWoody Haosheng Gan, William Barr Held, Diyi YangACL 2026
它引用的顶会 Paper23
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
相关 Paper
- SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM EvaluationGayathri Saranathan, Cong Xu, Mahammad Parwez Alam, Tarun Kumar 等ACL 2025 · 被引用 3 次
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
- Expanding the AI Evaluation Toolbox with Statistical ModelsDrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao 等ICML 2026 · 被引用 4 次
- Examining the robustness of LLM evaluation to the distributional assumptions of benchmarksCharlotte Siska, Katerina Marazopoulou, Melissa Ailem, James BonoACL 2024
- Learning More from Less: Unlocking Internal Representations for Benchmark CompressionYueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan 等ICML 2026
