Conformal Off-Policy Prediction in Contextual Bandits
Muhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh, Arnaud Doucet
摘要
Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture the variability of the outcome. In addition, particularly in safety-critical settings, stronger guarantees than asymptotic correctness may be required. To address these limitations, we consider a novel application of conformal prediction to contextual bandits. Given data collected under a behavioral policy, we propose conformal off-policy prediction (COPP), which can output reliable predictive intervals for the outcome under a new target policy. We provide theoretical finite-sample guarantees without making any additional assumptions beyond the standard contextual bandit setup, and empirically demonstrate the utility of COPP compared with existing methods on synthetic and real-world data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Conformal Meta-learners for Predictive Inference of Individual Treatment EffectsAhmed M. Alaa, Zaid Ahmad, Mark J. van der LaanNeurIPS 2023 · 被引用 32 次
- Conformal Prediction for Causal Effects of Continuous TreatmentsMaresa Schröder, Dennis Frauen, Jonas Schweisthal, Konstantin Hess 等NeurIPS 2025 · 被引用 21 次
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho 等ICML 2026 · 被引用 3 次
- Conformal Prediction for Verifiable Learned Query OptimizationHanwen Liu, Shashank Giridhara, Ibrahim SabekVLDB 2025 · 被引用 2 次
- Efficient Online Set-valued Classification with Bandit FeedbackZhou Wang, Xingye QiaoICML 2024 · 被引用 2 次
它引用的顶会 Paper6
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Learning Optimal Conformal ClassifiersDavid Stutz, Krishnamurthy Dvijotham, Ali Taylan Cemgil, Arnaud DoucetICLR 2022 · 被引用 123 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Off-Policy Risk Assessment in Contextual BanditsAudrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar AzizzadenesheliNeurIPS 2021 · 被引用 44 次
相关 Paper
- Off-Policy Confidence SequencesNikos Karampatziakis, Paul Mineiro, Aaditya RamdasICML 2021 · 被引用 19 次
- Conformal Bayesian ComputationEdwin Fong, Chris C. HolmesNeurIPS 2021 · 被引用 58 次
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- Counterfactual Learning with General Data-Generating PoliciesYusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei YataAAAI 2023 · 被引用 2 次
- Distribution-informed Online Conformal PredictionDongjian Hu, Junxi Wu, Shu-Tao Xia, Changliang ZouICLR 2026 · 被引用 2 次
