Counterfactual Learning with General Data-Generating Policies
Yusuke Narita, Kyohei Okumura, Akihiro Shimizu, Kohei Yata
Abstract
Off-policy evaluation (OPE) attempts to predict the performance of counterfactual policies using log data from a different policy. We extend its applicability by developing an OPE method for a class of both full support and deficient support logging policies in contextual-bandit settings. This class includes deterministic bandit (such as Upper Confidence Bound) as well as deterministic decision-making based on supervised and unsupervised learning. We prove that our method's prediction converges in probability to the true performance of a counterfactual policy as the sample size increases. We validate our method with experiments on partly and entirely deterministic logging policies. Finally, we apply it to evaluate coupon targeting policies by a major online platform and show how to improve the existing policy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cf20dcd-d937-40d0-9ae1-313dc562e6afCited by top-tier papers1
Ask how each one uses itBuilds on3
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 161 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 22 citations
Related papers
- Off-Policy Learning with Limited SupplyKoichi Tanaka, Ren Kishimoto, Bushun Kawagishi, Yusuke Narita et al.WWW 2026
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- Conformal Off-Policy Prediction in Contextual BanditsMuhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh et al.NeurIPS 2022 · 34 citations
- Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsHaanvid Lee, Jongmin Lee, Yunseon Choi, Wonseok Jeon et al.NeurIPS 2022 · 7 citations
