Empirical Likelihood for Contextual Bandits
Nikos Karampatziakis, John Langford, Paul Mineiro
Abstract
We propose an estimator and confidence interval for computing the value of a policy from off-policy data in the contextual bandit setting. To this end we apply empirical likelihood techniques to formulate our estimator and confidence interval as simple convex optimization problems. Using the lower bound of our confidence interval, we then propose an off-policy policy optimization algorithm that searches for policies with large reward lower bound. We empirically find that both our estimator and confidence interval improve over previous proposals in finite sample regimes. Finally, the policy optimization algorithm we propose outperforms a strong baseline system for learning from off-policy data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50dffb08-43fa-445f-8da6-ce92600cf467Cited by top-tier papers4
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li et al.NeurIPS 2020 · 96 citations
- On the Optimality of Batch Policy Optimization AlgorithmsChenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai et al.ICML 2021 · 36 citations
- Empirical Likelihood for Fair ClassificationPangpang Liu, Yichuan ZhaoICLR 2024 · 1 citation
- Understanding Deep Generative Models with Generalized Empirical LikelihoodsSuman V. Ravuri, Mélanie Rey, Shakir Mohamed, Marc Peter DeisenrothCVPR 2023
Related papers
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Off-Policy Interval Estimation with Lipschitz Value IterationZiyang Tang, Yihao Feng, Na Zhang, Jian Peng et al.NeurIPS 2020 · 6 citations
- Optimal Regret for Policy Optimization in Contextual BanditsOrin Levy, Yishay MansourICML 2026 · 1 citation
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 55 citations
