Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution Shifts
Guangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo Simeone
Abstract
We study the problem of monitoring model performance in dynamic environments where labeled data are limited. To this end, we propose prediction-powered risk monitoring (PPRM), a semi-supervised risk-monitoring approach based on prediction-powered inference (PPI). PPRM constructs anytime-valid lower bounds on the running risk by combining synthetic labels with a small set of true labels. Harmful shifts are detected via a threshold-based comparison with an upper bound on the nominal risk, satisfying assumption-free finite-sample guarantees on the type-I error. We demonstrate the effectiveness of PPRM through extensive experiments on image classification, large language model (LLM), and telecommunications monitoring tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c114ca90-b389-405a-a3b1-a8eedb5811bcBuilds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 36 citations
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 34 citations
- Protected Test-Time Adaptation via Online Entropy Matching: A Betting ApproachYarin Bar, Shalev Shaer, Yaniv RomanoNeurIPS 2024 · 27 citations
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra et al.NeurIPS 2024 · 27 citations
Related papers
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy et al.NeurIPS 2025 · 5 citations
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered InferencePranav Mani, Peng Xu, Zachary Lipton, Michael OberstICML 2026 · 8 citations
- Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc RegressionBenjamin Eyre, David MadrasICML 2025
- Semi-Supervised Hypothesis Testing by Betting on PredictionsYaniv Tenzer, Elad Tolochinksy, Yaniv RomanoICML 2026
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency GuaranteesSangwoo Park, Matteo Zecchin, Osvaldo SimeoneNeurIPS 2025 · 10 citations
