Empirical Likelihood for Contextual Bandits
Nikos Karampatziakis, John Langford, Paul Mineiro
摘要
We propose an estimator and confidence interval for computing the value of a policy from off-policy data in the contextual bandit setting. To this end we apply empirical likelihood techniques to formulate our estimator and confidence interval as simple convex optimization problems. Using the lower bound of our confidence interval, we then propose an off-policy policy optimization algorithm that searches for policies with large reward lower bound. We empirically find that both our estimator and confidence interval improve over previous proposals in finite sample regimes. Finally, the policy optimization algorithm we propose outperforms a strong baseline system for learning from off-policy data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li 等NeurIPS 2020 · 被引用 96 次
- On the Optimality of Batch Policy Optimization AlgorithmsChenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai 等ICML 2021 · 被引用 36 次
- Empirical Likelihood for Fair ClassificationPangpang Liu, Yichuan ZhaoICLR 2024 · 被引用 1 次
- Understanding Deep Generative Models with Generalized Empirical LikelihoodsSuman V. Ravuri, Mélanie Rey, Shakir Mohamed, Marc Peter DeisenrothCVPR 2023
相关 Paper
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 被引用 21 次
- Off-Policy Interval Estimation with Lipschitz Value IterationZiyang Tang, Yihao Feng, Na Zhang, Jian Peng 等NeurIPS 2020 · 被引用 6 次
- Optimal Regret for Policy Optimization in Contextual BanditsOrin Levy, Yishay MansourICML 2026 · 被引用 1 次
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 被引用 55 次
