No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
Pranav Mani, Peng Xu, Zachary Lipton, Michael Oberst
摘要
Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic free lunch for PPI++, an adaptive form of PPI, showing that the asymptotic variance of PPI++ is always less than or equal to the variance obtained from using gold-standard labels alone. Notably, this result holds regardless of the quality of the pseudo-labels. In this work, we demystify this result by conducting an exact finite-sample analysis of the estimation error of PPI++ on the mean estimation problem. We give a no free lunch result, characterizing the settings (and sample sizes) where PPI++ has provably worse estimation error than using gold-standard labels alone. Specifically, PPI++ will outperform if and only if the correlation between pseudo- and gold-standard is above a certain level that depends on the number of labeled samples (). In some cases our results simplify considerably: For Gaussian data, for instance, the correlation must be at least in order to see improvement. More broadly, by providing exact non-asymptotic expressions for the variance of PPI++ under sample splitting, we aim to empower practitioners to transparently reason about the benefits of PPI++ in specific applications. In experiments, we illustrate that our theoretical findings hold on real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
- Efficient Randomized Experiments Using Foundation ModelsPiersilvio De Bartolomeis, Javier Abad, Guanbo Wang, Konstantin Donhauser 等NeurIPS 2025 · 被引用 22 次
- Revisiting Active Sequential Prediction-Powered Mean EstimationMaria-Eleni Sfyraki, Jun-Kun WangICLR 2026 · 被引用 4 次
- Black-Box Assisted Regression: Phase Transitions and Minimax OptimalityYan ZhouICML 2026
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong 等ICLR 2026 · 被引用 150 次
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- Prediction-powered Generalization of Causal InferencesIlker Demirel, Ahmed M. Alaa, Anthony Philippakis, David A. SontagICML 2024 · 被引用 18 次
- Prediction-Powered Adaptive Shrinkage EstimationSida Li, Nikolaos IgnatiadisICML 2025
相关 Paper
- Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc RegressionBenjamin Eyre, David MadrasICML 2025
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy 等NeurIPS 2025 · 被引用 5 次
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra 等NeurIPS 2024 · 被引用 27 次
- FAB-PPI: Frequentist, Assisted by Bayes, Prediction-Powered InferenceStefano Cortinovis, Francois CaronICML 2025
- Prediction-Powered Adaptive Inference with Pretrained AI Models for Contextual BanditsGabriel Sargent, Wei Sun, Zhengwu Zhang, Yufeng LiuICML 2026
