Rethinking LLM Evaluation: Can We Evaluate LLMs with 200× Less Data?
Shaobo Wang, Cong Wang, Wenjie Fu, Yue Min, Mingquan Feng, Isabel Guan, Xuming Hu, Conghui He, Cunxiang Wang, Kexin Yang, Xingzhang Ren, Fei Huang
Abstract
Benchmark suites for large language models are growing faster than our ability to pay for them. Even when training is already expensive, many use cases require repeated evaluation across many checkpoints, variants, and competing systems, and the steady expansion of benchmark suites increasingly turns evaluation into a bottleneck in tokens and compute. This scale changes what ``useful data'' means. Instead of asking whether an instance is good for training one model, we ask which instances are necessary to keep the collective ordering of many models stable. We analyze redundancy at the instance level and find repetition in both the text and the ranking patterns induced across models. Based on this observation, we formulate benchmark compression as a subset optimization problem that targets accurate score reconstruction and ranking preservation at the same time. We propose EssenceBench, a coarse-to-fine framework with three stages: redundancy-aware filtering with text and ranking signals, fitness-driven subset search with an iterative genetic algorithm and a fixed surrogate predictor, and attribution-guided refinement for better coverage under tight budgets. Across multiple leaderboards, EssenceBench achieves lower reconstruction error and stronger ranking preservation than prior approaches while reducing selection time. On HellaSwag with 10K instances, EssenceBench preserves 95% of model rankings within a 5% shift using only 50 instances, a 200 compression. The source code will be made available upon acceptance of the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction TuningWei Liu, Weihao Zeng, Keqing He, Yong Jiang et al.ICLR 2024 · 369 citations
Related papers
- metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language ModelsAlexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric SchulzICLR 2025
- SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM EvaluationGayathri Saranathan, Cong Xu, Mahammad Parwez Alam, Tarun Kumar et al.ACL 2025 · 3 citations
- Learning More from Less: Unlocking Internal Representations for Benchmark CompressionYueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan et al.ICML 2026
- MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language ModelsZhongzhan Huang, Guoming Ling, Shanshan Zhong, Hefeng Wu et al.ACL 2025
- DataDecide: How to Predict Best Pretraining Data with Small ExperimentsIan Magnusson, Nguyen Tai, Ben Bogin, David Heineman et al.ICML 2025
