Off-Policy Confidence Sequences
Nikos Karampatziakis, Paul Mineiro, Aaditya Ramdas
2021年份
19被引次数
1顶会引用
摘要
We develop confidence bounds that hold uniformly over time for off-policy evaluation in the contextual bandit setting. These confidence sequences are based on recent ideas from martingale analysis and are non-asymptotic, non-parametric, and valid at arbitrary stopping times. We provide algorithms for computing these confidence sequences that strike a good balance between computational and statistical efficiency. We empirically demonstrate the tightness of our approach in terms of failure probability and width and apply it to the"gated deployment"problem of safely upgrading a production contextual bandit system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Conformal Off-Policy Prediction in Contextual BanditsMuhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh 等NeurIPS 2022 · 被引用 34 次
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- Confidence sequences for sampling without replacementIan Waudby-Smith, Aaditya RamdasNeurIPS 2020 · 被引用 57 次
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz 等NeurIPS 2021 · 被引用 58 次
