Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
Sangwoo Park, Matteo Zecchin, Osvaldo Simeone
摘要
Selecting artificial intelligence (AI) models, such as large language models (LLMs), from multiple candidates requires accurate performance estimation. This is ideally achieved through empirical evaluations involving abundant real-world data. However, such evaluations are costly and impractical at scale. To address this challenge, autoevaluation methods leverage synthetic data produced by automated evaluators, such as LLMs-as-judges, reducing variance but potentially introducing bias. Recent approaches have employed semi-supervised prediction-powered inference (PPI) to correct for the bias of autoevaluators. However, the use of autoevaluators may lead in practice to a degradation in sample efficiency compared to conventional methods using only real-world data. In this paper, we propose R-AutoEval+, a novel framework that provides finite-sample reliability guarantees on the model evaluation, while also ensuring an enhanced (or at least no worse) sample efficiency compared to conventional methods. The key innovation of R-AutoEval+ is an adaptive construction of the model evaluation variable, which dynamically tunes its reliance on synthetic data, reverting to conventional methods when the autoevaluator is insufficiently accurate. Experiments on the use of LLMs-as-judges for the optimization of quantization settings for the weights of an LLM, for prompt design in LLMs, and for test-time reasoning budget allocation in LLMs confirm the reliability and efficiency of R-AutoEval+.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Revisiting Active Sequential Prediction-Powered Mean EstimationMaria-Eleni Sfyraki, Jun-Kun WangICLR 2026 · 被引用 4 次
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI EvaluationYizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi WangICML 2026 · 被引用 1 次
- Margin-Adaptive Confidence Ranking for Reliable LLM JudgementGaojie Jin, Yong Tao, Lijia Yu, Tianjin HuangICML 2026
- Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution ShiftsGuangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo SimeoneICML 2026
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster 等ICLR 2023 · 被引用 297 次
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong 等ICLR 2026 · 被引用 150 次
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 被引用 48 次
- Active, anytime-valid risk controlling prediction setsZiyu Xu, Nikos Karampatziakis, Paul MineiroNeurIPS 2024 · 被引用 19 次
相关 Paper
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo 等ICML 2026 · 被引用 5 次
- Routing and Reasoned Evaluation with Large Language ModelsGuiyao Tie, Tianyao Luo, Xueyang Zhou, Chaoran Hu 等ICML 2026
- Multiple-Prediction-Powered InferenceCharlie Cowen-Breen, Alekh Agarwal, Stephen Bates, William W. Cohen 等ICLR 2026 · 被引用 11 次
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn 等ICML 2026 · 被引用 24 次
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li 等AAAI 2026
