Semi-Supervised Hypothesis Testing by Betting on Predictions
Yaniv Tenzer, Elad Tolochinksy, Yaniv Romano
摘要
We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of , and additional unlabeled samples from the marginal of , we ask how unlabeled data can be used to hypothesize about the distribution of , and the conditional distribution of . We introduce an e-statistic and use it to construct a sequential test. Under standard distributional assumptions---label shift or concept shift---we establish that the test is anytime valid. Furthermore, we show that for binary data, the e-statistic has non-trivial power. Crucially, our approach retains these properties even when the underlying predictions are inaccurate. Through simulations and applications to large language models evaluation, we demonstrate power gains over baseline approaches, including prediction-powered inference. These gains persist even with relatively limited unlabeled data and when predictions have low accuracy due to weak correlation between and .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
- Anytime-valid, Bayes-assisted, Prediction-Powered InferenceValentin Kilian, Stefano Cortinovis, Francois CaronNeurIPS 2025 · 被引用 9 次
- Prediction-Powered E-ValuesDaniel Csillag, Cláudio José Struchiner, Guilherme Tegoni GoedertICML 2025
相关 Paper
- Sequential Predictive Two-Sample and Independence TestingAleksandr Podkopaev, Aaditya RamdasNeurIPS 2023 · 被引用 29 次
- Auditing Fairness by BettingBen Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya RamdasNeurIPS 2023 · 被引用 29 次
- Sequential Kernel Goodness-of-fit TestingZhengyu Zhou, Weiwei LiuICML 2024 · 被引用 1 次
- Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution ShiftsGuangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo SimeoneICML 2026
- Prediction-Powered Adaptive Inference with Pretrained AI Models for Contextual BanditsGabriel Sargent, Wei Sun, Zhengwu Zhang, Yufeng LiuICML 2026
