Conformal Off-Policy Prediction in Contextual Bandits
Muhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh, Arnaud Doucet
Abstract
Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture the variability of the outcome. In addition, particularly in safety-critical settings, stronger guarantees than asymptotic correctness may be required. To address these limitations, we consider a novel application of conformal prediction to contextual bandits. Given data collected under a behavioral policy, we propose conformal off-policy prediction (COPP), which can output reliable predictive intervals for the outcome under a new target policy. We provide theoretical finite-sample guarantees without making any additional assumptions beyond the standard contextual bandit setup, and empirically demonstrate the utility of COPP compared with existing methods on synthetic and real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6597e7d2-5863-49a3-9bac-a827b76c9fa7Cited by top-tier papers8
- Conformal Meta-learners for Predictive Inference of Individual Treatment EffectsAhmed M. Alaa, Zaid Ahmad, Mark J. van der LaanNeurIPS 2023 · 32 citations
- Conformal Prediction for Causal Effects of Continuous TreatmentsMaresa Schröder, Dennis Frauen, Jonas Schweisthal, Konstantin Hess et al.NeurIPS 2025 · 21 citations
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho et al.ICML 2026 · 3 citations
- Conformal Prediction for Verifiable Learned Query OptimizationHanwen Liu, Shashank Giridhara, Ibrahim SabekVLDB 2025 · 2 citations
- Efficient Online Set-valued Classification with Bandit FeedbackZhou Wang, Xingye QiaoICML 2024 · 2 citations
Builds on6
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Learning Optimal Conformal ClassifiersDavid Stutz, Krishnamurthy Dvijotham, Ali Taylan Cemgil, Arnaud DoucetICLR 2022 · 123 citations
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 86 citations
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller et al.NeurIPS 2021 · 64 citations
- Off-Policy Risk Assessment in Contextual BanditsAudrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar AzizzadenesheliNeurIPS 2021 · 44 citations
Related papers
- Off-Policy Confidence SequencesNikos Karampatziakis, Paul Mineiro, Aaditya RamdasICML 2021 · 19 citations
- Conformal Bayesian ComputationEdwin Fong, Chris C. HolmesNeurIPS 2021 · 58 citations
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
- Counterfactual Learning with General Data-Generating PoliciesYusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei YataAAAI 2023 · 2 citations
- Distribution-informed Online Conformal PredictionDongjian Hu, Junxi Wu, Shu-Tao Xia, Changliang ZouICLR 2026 · 2 citations
