Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression
Benjamin Eyre, David Madras
Abstract
The availability of machine learning systems that can effectively perform arbitrary tasks has led to synthetic labels from these systems being used in applications of statistical inference, such as data analysis or model evaluation. The Prediction Powered Inference (PPI) framework provides a way of leveraging both a large pool of pseudo-labelled data and a small sample with real, high-quality labels to produce a low-variance, unbiased estimate of the quantity being evaluated for. Most work on PPI considers a relatively sizable set of labelled samples, which can be resource intensive to obtain. However, we find that when labelled data is scarce, the PPI++ method can perform even worse than classical inference. We analyze this phenomenon by relating PPI++ to ordinary least squares regression, which also experiences high variance with small sample sizes, and use this regression framework to better understand the efficacy of PPI. Motivated by this, we present two new PPI-based techniques that leverage robust regressors to produce even lower variance estimators in the few-label regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac81b602-7b17-4a38-8e23-3e4c238e2cd0Cited by top-tier papers3
- Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect JudgesChen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle et al.ICLR 2026 · 24 citations
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered InferencePranav Mani, Peng Xu, Zachary Lipton, Michael OberstICML 2026 · 8 citations
- Revisiting Active Sequential Prediction-Powered Mean EstimationMaria-Eleni Sfyraki, Jun-Kun WangICLR 2026 · 4 citations
Builds on5
- Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal TrainingYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ACL 2025 · 65 citations
- MoCoDA: Model-based Counterfactual Data AugmentationSilviu Pitis, Elliot Creager, Ajay Mandlekar, Animesh GargNeurIPS 2022 · 60 citations
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 34 citations
- When can Regression-Adjusted Control Variate Help? Rare Events, Sobolev Embedding and Minimax OptimalityJose H. Blanchet, Haoxuan Chen, Yiping Lu, Lexing YingNeurIPS 2023 · 6 citations
- AutoEval Done Right: Using Synthetic Data for Model EvaluationPierre Boyeau, Anastasios Nikolas Angelopoulos, Tianle Li, Nir Yosef et al.ICML 2025
Related papers
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy et al.NeurIPS 2025 · 5 citations
- MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean EstimationSe Yoon Lee, Jae-kwang KimICML 2026 · 1 citation
- FAB-PPI: Frequentist, Assisted by Bayes, Prediction-Powered InferenceStefano Cortinovis, Francois CaronICML 2025
- Prediction-Powered Adaptive Shrinkage EstimationSida Li, Nikolaos IgnatiadisICML 2025
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra et al.NeurIPS 2024 · 27 citations
