Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning
Alberto Maria Metelli, Alessio Russo, Marcello Restelli
Abstract
Importance Sampling (IS) is a widely used building block for a large variety of off-policy estimation and learning algorithms. However, empirical and theoretical studies have progressively shown that vanilla IS leads to poor estimations whenever the behavioral and target policies are too dissimilar. In this paper, we analyze the theoretical properties of the IS estimator by deriving a probabilistic deviation lower bound that formalizes the intuition behind its undesired behavior. Then, we propose a class of IS transformations, based on the notion of power mean, that are able, under certain circumstances, to achieve a subgaussian concentration rate. Differently from existing methods, like weight truncation, our estimator preserves the differentiability in the target distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cbbf509-3da9-474e-be21-87c7f738c160Cited by top-tier papers22
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 34 citations
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Off-Policy Evaluation with Deficient Support Using Side InformationNicolò Felicioni, Maurizio Ferrari Dacrema, Marcello Restelli, Paolo CremonesiNeurIPS 2022 · 19 citations
Builds on2
Related papers
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- BR-SNIS: Bias Reduced Self-Normalized Importance SamplingGabriel Cardoso, Sergey Samsonov, Achille Thin, Eric Moulines et al.NeurIPS 2022 · 21 citations
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 49 citations
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
- SOPE: Spectrum of Off-Policy EstimatorsChristina J. Yuan, Yash Chandak, Stephen Giguere, Philip S. Thomas et al.NeurIPS 2021 · 6 citations
