Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and Learning
Alberto Maria Metelli, Alessio Russo, Marcello Restelli
摘要
Importance Sampling (IS) is a widely used building block for a large variety of off-policy estimation and learning algorithms. However, empirical and theoretical studies have progressively shown that vanilla IS leads to poor estimations whenever the behavioral and target policies are too dissimilar. In this paper, we analyze the theoretical properties of the IS estimator by deriving a probabilistic deviation lower bound that formalizes the intuition behind its undesired behavior. Then, we propose a class of IS transformations, based on the notion of power mean, that are able, under certain circumstances, to achieve a subgaussian concentration rate. Differently from existing methods, like weight truncation, our estimator preserves the differentiability in the target distribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito 等AAAI 2023 · 被引用 29 次
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 被引用 21 次
- Off-Policy Evaluation with Deficient Support Using Side InformationNicolò Felicioni, Maurizio Ferrari Dacrema, Marcello Restelli, Paolo CremonesiNeurIPS 2022 · 被引用 19 次
它引用的顶会 Paper2
相关 Paper
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- BR-SNIS: Bias Reduced Self-Normalized Importance SamplingGabriel Cardoso, Sergey Samsonov, Achille Thin, Eric Moulines 等NeurIPS 2022 · 被引用 21 次
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 被引用 49 次
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht 等NeurIPS 2022 · 被引用 19 次
- SOPE: Spectrum of Off-Policy EstimatorsChristina J. Yuan, Yash Chandak, Stephen Giguere, Philip S. Thomas 等NeurIPS 2021 · 被引用 6 次
