Offline Policy Evaluation in Large Action Spaces via Outcome-Oriented Action Grouping
Jie Peng, Hao Zou, Jiashuo Liu, Shaoming Li, Yibao Jiang, Jian Pei, Peng Cui
摘要
Offline policy evaluation (OPE) aims to accurately estimate the performance of a hypothetical policy using only historical data, which has drawn increasing attention in a wide range of applications including recommender systems and personalized medicine. With the presence of rising granularity of consumer data, many industries started exploring larger action candidate spaces to support more precise personalized action. While inverse propensity score (IPS) is a standard OPE estimator, it suffers from more severe variance issues with increasing action spaces. To address this issue, we theoretically prove that the estimation variance can be reduced by merging actions into groups while the distinction among these action effects on the outcome can induce extra bias. Motivated by these, we propose a novel IPS estimator with outcome-oriented action Grouping (GroupIPS), which leverages a Lipschitz regularized network to measure the distance of action effects in the embedding space and merges nearest action neighbors. This strategy enables more robust estimation by achieving smaller variances while inducing minor additional bias. Empirically, extensive experiments on both synthetic and real world datasets demonstrate the effectiveness of our proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper11
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 被引用 21 次
- Off-Policy Evaluation for Large Action Spaces via Policy ConvolutionNoveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus 等WWW 2024 · 被引用 17 次
- Long-term Off-Policy Evaluation and LearningYuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette 等WWW 2024 · 被引用 15 次
- Off-policy estimation with adaptively collected data: the power of online learningJeonghwan Lee, Cong MaNeurIPS 2024 · 被引用 4 次
相关 Paper
- Off-Policy Evaluation of Ranking Policies under Diverse User BehaviorHaruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu 等KDD 2023 · 被引用 8 次
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 被引用 18 次
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Uncertainty-Aware Instance Reweighting for Off-Policy LearningXiaoying Zhang, Junpu Chen, Hongning Wang, Hong Xie 等NeurIPS 2023 · 被引用 6 次
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 被引用 14 次
