Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution Shifts
Guangyi Zhang, Yunlong Cai, Guanding Yu, Osvaldo Simeone
摘要
We study the problem of monitoring model performance in dynamic environments where labeled data are limited. To this end, we propose prediction-powered risk monitoring (PPRM), a semi-supervised risk-monitoring approach based on prediction-powered inference (PPI). PPRM constructs anytime-valid lower bounds on the running risk by combining synthetic labels with a small set of true labels. Harmful shifts are detected via a threshold-based comparison with an upper bound on the nominal risk, satisfying assumption-free finite-sample guarantees on the type-I error. We demonstrate the effectiveness of PPRM through extensive experiments on image classification, large language model (LLM), and telecommunications monitoring tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tracking the risk of a deployed model and detecting harmful distribution shiftsAleksandr Podkopaev, Aaditya RamdasICLR 2022 · 被引用 36 次
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- Protected Test-Time Adaptation via Online Entropy Matching: A Betting ApproachYarin Bar, Shalev Shaer, Yaniv RomanoNeurIPS 2024 · 被引用 27 次
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra 等NeurIPS 2024 · 被引用 27 次
相关 Paper
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy 等NeurIPS 2025 · 被引用 5 次
- No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered InferencePranav Mani, Peng Xu, Zachary Lipton, Michael OberstICML 2026 · 被引用 8 次
- Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc RegressionBenjamin Eyre, David MadrasICML 2025
- Semi-Supervised Hypothesis Testing by Betting on PredictionsYaniv Tenzer, Elad Tolochinksy, Yaniv RomanoICML 2026
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency GuaranteesSangwoo Park, Matteo Zecchin, Osvaldo SimeoneNeurIPS 2025 · 被引用 10 次
