Multiple-Prediction-Powered Inference
Charlie Cowen-Breen, Alekh Agarwal, Stephen Bates, William W. Cohen, Jacob Eisenstein, Amir Globerson, Adam Fisch
摘要
Statistical estimation often involves tradeoffs between expensive, high-quality measurements and a variety of lower-quality proxies. We introduce Multiple-Prediction-Powered Inference (MultiPPI): a general framework for constructing statistically efficient estimates by optimally allocating resources across these diverse data sources. This work provides theoretical guarantees about the minimax optimality, finite-sample performance, and asymptotic normality of the MultiPPI estimator, and through experiments across three diverse large language model (LLM) evaluation scenarios, we show that MultiPPI consistently achieves lower estimation error than existing baselines. This advantage stems from its budget-adaptive allocation strategy, which strategically combines subsets of models by learning their complex cost and correlation structures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin 等ACL 2025 · 被引用 209 次
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 被引用 81 次
- Prediction-Powered Ranking of Large Language ModelsIvi Chatzi, Eleni Straitouri, Suhas Thejaswi, Manuel Gomez RodriguezNeurIPS 2024 · 被引用 34 次
相关 Paper
- Prediction-Powered Adaptive Shrinkage EstimationSida Li, Nikolaos IgnatiadisICML 2025
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency GuaranteesSangwoo Park, Matteo Zecchin, Osvaldo SimeoneNeurIPS 2025 · 被引用 10 次
- Efficient Inference for Noisy LLM-as-a-Judge EvaluationYiqun Chen, Sizhu Lu, Sijia Li, Moran Guo 等ICML 2026 · 被引用 5 次
- Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc RegressionBenjamin Eyre, David MadrasICML 2025
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered InferencePranav Mani, Peng Xu, Zachary Lipton, Michael OberstICML 2026 · 被引用 8 次
