POTEC: Off-Policy Contextual Bandits for Large Action Spaces via Policy Decomposition
Yuta Saito, Jihan Yao, Thorsten Joachims
摘要
We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -most of which rely crucially on reward-regression models or importanceweighted policy gradients -fail due to excessive bias or variance. To overcome these issues in OPL, we propose a novel two-stage algorithm, called Policy Optimization via Two-Stage Policy Decomposition (POTEC). It leverages clustering in the action space and learns two different policies via policy-and regression-based approaches, respectively. In particular, we derive a novel lowvariance gradient estimator that enables to learn a first-stage policy for cluster selection efficiently via a policy-based approach. To select a specific action within the cluster sampled by the first-stage policy, POTEC uses a second-stage policy derived from a regression-based approach within each cluster. We show that a local correctness condition, which only requires that the regression model preserves the relative expected reward differences of the actions within each cluster, ensures that our policy-gradient estimator is unbiased and the second-stage policy is optimal. We also show that POTEC provides a strict generalization of policy-and regression-based approaches and their associated assumptions. Comprehensive experiments demonstrate that POTEC provides substantial improvements in OPL effectiveness particularly in large and structured action spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningAlberto Maria Metelli, Alessio Russo, Marcello RestelliNeurIPS 2021 · 被引用 55 次
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 被引用 55 次
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 被引用 49 次
相关 Paper
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- Local Clustering in Contextual Multi-Armed BanditsYikun Ban, Jingrui HeWWW 2021 · 被引用 51 次
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 被引用 35 次
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- Off-Policy Evaluation for Large Action Spaces via Policy ConvolutionNoveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus 等WWW 2024 · 被引用 17 次
