On Speeding Up Language Model Evaluation
Jin Peng Zhou, Christian K. Belardi, Ruihan Wu, Travis Zhang, Carla P. Gomes, Wen Sun, Kilian Q. Weinberger
摘要
Developing prompt-based methods with Large Language Models (LLMs) requires making numerous decisions, which give rise to a combinatorial search problem over hyper-parameters. This exhaustive evaluation can be time-consuming and costly. In this paper, we propose an adaptive approach to explore this space. We are exploiting the fact that often only few samples are needed to identify clearly superior or inferior settings, and that many evaluation tests are highly correlated. We lean on multiarmed bandits to sequentially identify the next (method, validation sample)-pair to evaluate and utilize low-rank matrix factorization to fill in missing evaluations. We carefully assess the efficacy of our approach on several competitive benchmark problems and show that it can identify the top-performing method using only 5-15% of the typical resources-resulting in 85-95% LLM cost savings. Our code is available at https://github.com/kilian-group/banditeval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
- Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm That Provably Exploits Model SimilarityZifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang 等ICML 2026
- UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective OptimizationPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang 等ICLR 2025
- Accelerating Unbiased LLM Evaluation via Synthetic FeedbackZhaoyi Zhou, Yuda Song, Andrea ZanetteICML 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Efficient Prompt Optimization Through the Lens of Best Arm IdentificationChengshuai Shi, Kun Yang, Zihan Chen, Jundong Li 等NeurIPS 2024 · 被引用 44 次
- Efficient Multi-objective Prompt Optimization via Pure-exploration BanditsDonghao Li, Chengshuai Shi, Weijuan Ou, Cong Shen 等ICLR 2026 · 被引用 2 次
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
- MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural NetworksZhi Hong, Qian Zhang, Jiahang Sun, Zhiwei Shang 等ICML 2026 · 被引用 3 次
- Online Multi-LLM Selection via Contextual Bandits Under Unstructured Context EvolutionManhin Poon, Xiangxiang Dai, Xutong Liu, Fang Kong 等AAAI 2026 · 被引用 11 次
