Off-Policy Evaluation for Large Action Spaces via Policy Convolution
Noveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus, Julian J. McAuley
摘要
Developing accurate off-policy estimators is crucial for both evaluating and optimizing for new policies. The main challenge in off-policy estimation is the distribution shift between the logging policy that generates data and the target policy that we aim to evaluate. Typically, techniques for correcting distribution shift involve some form of importance sampling. This approach results in unbiased value estimation but often comes with the trade-off of high variance, even in the simpler case of one-step contextual bandits. Furthermore, importance sampling relies on the common support assumption, which becomes impractical when the action space is large. To address these challenges, we introduce the Policy Convolution (PC) family of estimators. These methods leverage latent structure within actions-made available through action embeddings-to strategically convolve the logging and target policies. This convolution introduces a unique bias-variance trade-off, which can be controlled by adjusting the amount of convolution. Our experiments on synthetic and benchmark datasets demonstrate remarkable mean squared error (MSE) improvements when using PC, especially when either the action space or policy mismatch becomes large, with gains of up to 5 -6 orders of magnitude over existing estimators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 被引用 21 次
- Long-term Off-Policy Evaluation and LearningYuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette 等WWW 2024 · 被引用 15 次
- Exploiting Similarities in A/B Testing with Off-Policy EstimationOtmane Sakhi, Alexandre Gilotte, David RohdeKDD 2026 · 被引用 2 次
- Log-Sum-Exponential Estimator for Off-Policy Evaluation and LearningArmin Behnamnia, Gholamali Aminian, Alireza Aghaei, Chengchun Shi 等ICML 2025
- Doubly Robust Fusion of Many Treatments for Policy LearningKe Zhu, Jianing Chu, Ilya Lipkovich, Wenyu Ye 等ICML 2025
它引用的顶会 Paper19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 被引用 150 次
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 被引用 81 次
相关 Paper
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State AbstractionBrahma S. Pavse, Josiah P. HannaAAAI 2023 · 被引用 9 次
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- POTEC: Off-Policy Contextual Bandits for Large Action Spaces via Policy DecompositionYuta Saito, Jihan Yao, Thorsten JoachimsICLR 2025
