Off-Policy Evaluation for Large Action Spaces via Policy Convolution
Noveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus, Julian J. McAuley
Abstract
Developing accurate off-policy estimators is crucial for both evaluating and optimizing for new policies. The main challenge in off-policy estimation is the distribution shift between the logging policy that generates data and the target policy that we aim to evaluate. Typically, techniques for correcting distribution shift involve some form of importance sampling. This approach results in unbiased value estimation but often comes with the trade-off of high variance, even in the simpler case of one-step contextual bandits. Furthermore, importance sampling relies on the common support assumption, which becomes impractical when the action space is large. To address these challenges, we introduce the Policy Convolution (PC) family of estimators. These methods leverage latent structure within actions-made available through action embeddings-to strategically convolve the logging and target policies. This convolution introduces a unique bias-variance trade-off, which can be controlled by adjusting the amount of convolution. Our experiments on synthetic and benchmark datasets demonstrate remarkable mean squared error (MSE) improvements when using PC, especially when either the action space or policy mismatch becomes large, with gains of up to 5 -6 orders of magnitude over existing estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Long-term Off-Policy Evaluation and LearningYuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette et al.WWW 2024 · 15 citations
- Exploiting Similarities in A/B Testing with Off-Policy EstimationOtmane Sakhi, Alexandre Gilotte, David RohdeKDD 2026 · 2 citations
- Log-Sum-Exponential Estimator for Off-Policy Evaluation and LearningArmin Behnamnia, Gholamali Aminian, Alireza Aghaei, Chengchun Shi et al.ICML 2025
- Doubly Robust Fusion of Many Treatments for Policy LearningKe Zhu, Jianing Chu, Ilya Lipkovich, Wenyu Ye et al.ICML 2025
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Discretizing Continuous Action Space for On-Policy OptimizationYunhao Tang, Shipra AgrawalAAAI 2020 · 150 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang et al.WWW 2020 · 106 citations
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 81 citations
Related papers
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State AbstractionBrahma S. Pavse, Josiah P. HannaAAAI 2023 · 9 citations
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 34 citations
- POTEC: Off-Policy Contextual Bandits for Large Action Spaces via Policy DecompositionYuta Saito, Jihan Yao, Thorsten JoachimsICLR 2025
