Off-Policy Evaluation for Large Action Spaces via Embeddings
Yuta Saito, Thorsten Joachims
Abstract
Off-policy evaluation (OPE) in contextual bandits has seen rapid adoption in real-world systems, since it enables offline evaluation of new policies using only historic log data. Unfortunately, when the number of actions is large, existing OPE estimators -- most of which are based on inverse propensity score weighting -- degrade severely and can suffer from extreme bias and variance. This foils the use of OPE in many applications from recommender systems to language models. To overcome this issue, we propose a new OPE estimator that leverages marginalized importance weights when action embeddings provide structure in the action space. We characterize the bias, variance, and mean squared error of the proposed estimator and analyze the conditions under which the action embedding provides statistical benefits over conventional estimators. In addition to the theoretical analysis, we find that the empirical performance improvement can be substantial, enabling reliable OPE even when existing estimators collapse due to a large number of actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10d81851-e5a7-4d01-8ddd-a7ff5ff3b1eeCited by top-tier papers24
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez et al.NeurIPS 2022 · 63 citations
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 34 citations
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 18 citations
Builds on9
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningAlberto Maria Metelli, Alessio Russo, Marcello RestelliNeurIPS 2021 · 55 citations
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 55 citations
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 49 citations
- Counterfactual Evaluation of Slate Recommendations with Sequential Reward InteractionsJames McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra et al.KDD 2020 · 45 citations
Related papers
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 14 citations
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Off-Policy Evaluation for Large Action Spaces via Policy ConvolutionNoveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus et al.WWW 2024 · 17 citations
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 22 citations
- Offline Policy Evaluation in Large Action Spaces via Outcome-Oriented Action GroupingJie Peng, Hao Zou, Jiashuo Liu, Shaoming Li et al.WWW 2023 · 22 citations
