ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Yizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi Wang
摘要
Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose ProEval, a proactive evaluation framework that leverages transfer learning to efficiently estimate performance and identify failure cases. ProEval employs pre-trained Gaussian Processes (GPs) as surrogates for the performance score function, mapping model inputs to metrics such as the severity of errors or safety violations. By framing performance estimation as Bayesian quadrature (BQ) and failure discovery as superlevel set sampling, we develop uncertainty-aware decision strategies that actively select or synthesize highly informative inputs for testing. Theoretically, we prove that our pre-trained GP-based BQ estimator is unbiased and bounded. Empirically, extensive experiments on reasoning, safety alignment, and classification benchmarks demonstrate that ProEval is significantly more efficient than competitive baselines. It requires 8–65x fewer samples to achieve estimates within of the ground truth, while simultaneously revealing more diverse failure cases under a stricter evaluation budget. Our open-sourced code and data can be found at https://github.com/google-deepmind/proeval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster 等ICLR 2023 · 被引用 297 次
相关 Paper
- Gaussian Process Probes (GPP) for Uncertainty-Aware ProbingZi Wang, Alexander Ku, Jason Baldridge, Tom Griffiths 等NeurIPS 2023 · 被引用 17 次
- ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue AgentsTianjian Liu, Fanqi Wan, Jiajian Guo, Xiaojun QuanACL 2026
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr 等ICLR 2020 · 被引用 104 次
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionZhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu 等NeurIPS 2024 · 被引用 125 次
- Active Bayesian Assessment of Black-Box ClassifiersDisi Ji, Robert L. Logan IV, Padhraic Smyth, Mark SteyversAAAI 2021 · 被引用 3 次
