Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm That Provably Exploits Model Similarity
Zifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang, Florian Dorner
摘要
Large Language Models are typically benchmarked by evaluating every model on every test query. For practitioners seeking the best model to deploy, this is often wasteful: if a model clearly performs worse than others, there is no need to precisely estimate its performance. Best-arm identification algorithms can be naturally applied to drastically reduce costs by adaptively allocating evaluation budget. Further, language models often respond similarly to the same prompt—a property previous work has tried to leverage with mixed success in different use cases. We propose Synchronized Successive Rejects (SySRs), augmenting the classical Successive Reject algorithm with paired comparisons. Unlike prior attempts to leverage model similarity in best-model identification, our approach is hyperparameter-free and enjoys performance guarantees that improve with the degree of similarity between evaluated models. Empirically, our method outperforms all baselines in terms of average error rate across 15 standard benchmarks, and in terms of worst-case budget for reliably identifying the best model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- Accuracy on the Line: on the Strong Correlation Between Out-of-Distribution and In-Distribution GeneralizationJohn Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa 等ICML 2021 · 被引用 323 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 被引用 7 次
相关 Paper
- On Speeding Up Language Model EvaluationJin Peng Zhou, Christian K. Belardi, Ruihan Wu, Travis Zhang 等ICLR 2025
- On Universally Optimal Algorithms for A/B TestingPo-An Wang, Kaito Ariu, Alexandre ProutièreICML 2024 · 被引用 4 次
- Efficient Prompt Optimization Through the Lens of Best Arm IdentificationChengshuai Shi, Kun Yang, Zihan Chen, Jundong Li 等NeurIPS 2024 · 被引用 44 次
- Best Arm Identification with Fixed Budget: A Large Deviation PerspectivePo-An Wang, Ruo-Chun Tzeng, Alexandre ProutièreNeurIPS 2023 · 被引用 15 次
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
