Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation
Peiwen Yuan, Yueqi Zhang, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Jiayi Shi, Chuyi Tan, Boyuan Pan, Yao Hu, Kan Li
Abstract
Evaluating models on large benchmarks is very resource-intensive, especially during the period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them only on a small and static coreset of the benchmark, which is derived from the publicly available evaluation results of source models. These methods rely on the assumption that target models have high prediction consistency with source models. However, we demonstrate that it doesn't generalize well in practice. To alleviate the inconsistency issue, we present TAILOREDBENCH, a method that conducts customized evaluation tailored to each target model. Specifically, a Global-coreset is first constructed as a probe to identify the most consistent source models for each target model with an adaptive source model selection strategy. Afterwards, a scalable K-Medoids clustering algorithm is proposed to extend the Globalcoreset to a tailored Native-coreset for each target model. According to the predictions on Native-coresets, we obtain the performance of target models on the whole benchmark with a calibrated estimation strategy. Comprehensive experiments on 5 benchmarks across over 300 models demonstrate that compared to best performing baselines, TAILOREDBENCH achieves an average reduction of 31.4% in MAE of accuracy estimates under the same inference budgets, showcasing strong effectiveness and generalizability 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1563cc6-3750-4fb4-a60f-c14c2ee1aa39Cited by top-tier papers3
- SparseEval: Efficient Evaluation of Large Language Models by Sparse OptimizationTaolin Zhang, Hang Guo, Wang Lu, Tao Dai et al.ICLR 2026 · 6 citations
- DISCO: Diversifying Sample Condensation for Efficient Model EvaluationAlexander Rubinstein, Benjamin Raible, Martin Gubri, Seong Joon OhICLR 2026 · 4 citations
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
Builds on4
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- Predicting Emergent Abilities with Infinite Resolution EvaluationShengding Hu, Xin Liu, Xu Han, Xinrong Zhang et al.ICLR 2024 · 27 citations
- Predicting the Performance of Foundation Models via Agreement-on-the-LineRahul Saxena, Taeyoun Kim, Aman Mehra, Christina Baek et al.NeurIPS 2024 · 8 citations
Related papers
- Learning More from Less: Unlocking Internal Representations for Benchmark CompressionYueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan et al.ICML 2026
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba et al.ICML 2025
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
- ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended CapabilitiesAdhiraj Ghosh, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao et al.ACL 2025
