Query-efficient model evaluation using cached responses
Hayden Helm, Ben Johnson, Carey Priebe
摘要
Evaluating a new model on an existing benchmark is often necessary to understand its behavior before deployment. For modern evaluation frameworks, generating and evaluating a response for all queries can be prohibitively expensive. In practice, responses from previously-evaluated models are often cached -- creating a potential opportunity to use this additional information to decrease the number of queries required to accurately evaluate a new model. In this paper, we introduce an approach for predicting benchmark performance that leverages cached model responses based on the Data Kernel Perspective Space (DKPS), a method for quantifying the relationship between models in the black-box setting. Theoretically, we show that DKPS-based methods are query-efficient under certain conditions. Empirically, we demonstrate that DKPS-based methods achieve the same mean absolute error as baselines with a substantially decreased query budget. We conclude by proposing an offline method for selecting a set of queries that maximizes the goodness-of-fit on reference models, improving prediction accuracy over random query selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- Tracking the perspectives of interacting language modelsHayden S. Helm, Brandon Duderstadt, Youngser Park, Carey E. PriebeEMNLP 2024
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
相关 Paper
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
- KernelBench: Can LLMs Write Efficient GPU Kernels?Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang 等ICML 2025
- Learning More from Less: Unlocking Internal Representations for Benchmark CompressionYueqi Zhang, Jin Hu, Shaoxiong Feng, Peiwen Yuan 等ICML 2026
- Detecting Perspective Shifts in Multi-Agent SystemsEric Bridgeford, Hayden HelmICML 2026 · 被引用 4 次
- Hyperband-based Bayesian Optimization for Black-box Prompt SelectionLennart Schneider, Martin Wistuba, Aaron Klein, Jacek Golebiowski 等ICML 2025
