Semi-Supervised Hypothesis Testing by Betting on Predictions
Yaniv Tenzer, Elad Tolochinksy, Yaniv Romano
Abstract
We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of , and additional unlabeled samples from the marginal of , we ask how unlabeled data can be used to hypothesize about the distribution of , and the conditional distribution of . We introduce an e-statistic and use it to construct a sequential test. Under standard distributional assumptions---label shift or concept shift---we establish that the test is anytime valid. Furthermore, we show that for binary data, the e-statistic has non-trivial power. Crucially, our approach retains these properties even when the underlying predictions are inaccurate. Through simulations and applications to large language models evaluation, we demonstrate power gains over baseline approaches, including prediction-powered inference. These gains persist even with relatively limited unlabeled data and when predictions have low accuracy due to weak correlation between and .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 96 citations
- Anytime-valid, Bayes-assisted, Prediction-Powered InferenceValentin Kilian, Stefano Cortinovis, Francois CaronNeurIPS 2025 · 9 citations
- Prediction-Powered E-ValuesDaniel Csillag, Cláudio José Struchiner, Guilherme Tegoni GoedertICML 2025
Related papers
- Sequential Predictive Two-Sample and Independence TestingAleksandr Podkopaev, Aaditya RamdasNeurIPS 2023 · 29 citations
- Auditing Fairness by BettingBen Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya RamdasNeurIPS 2023 · 29 citations
- Sequential Kernel Goodness-of-fit TestingZhengyu Zhou, Weiwei LiuICML 2024 · 1 citation
- Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution ShiftsGuangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo SimeoneICML 2026
- Prediction-Powered Adaptive Inference with Pretrained AI Models for Contextual BanditsGabriel Sargent, Wei Sun, Zhengwu Zhang, Yufeng LiuICML 2026
