State Relevance for Off-Policy Evaluation
Simon P. Shen, Yecheng Jason Ma, Omer Gottesman, Finale Doshi-Velez
摘要
Importance sampling-based estimators for off-policy evaluation (OPE) are valued for their simplicity, unbiasedness, and reliance on relatively few assumptions. However, the variance of these estimators is often high, especially when trajectories are of different lengths. In this work, we introduce Omitting-States-Irrelevant-to-Return Importance Sampling (OSIRIS), an estimator which reduces variance by strategically omitting likelihood ratios associated with certain states. We formalize the conditions under which OSIRIS is unbiased and has lower variance than ordinary importance sampling, and we demonstrate these properties empirically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Conservative Offline Distributional Reinforcement LearningYecheng Jason Ma, Dinesh Jayaraman, Osbert BastaniNeurIPS 2021 · 被引用 118 次
- Smooth Multi-Policy Causal Effect Estimation in Longitudinal SettingsWenxin Chen, Weishen Pan, Kyra Gan, Fei WangICML 2026
它引用的顶会 Paper2
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential TransitionsOmer Gottesman, Joseph Futoma, Yao Liu, Sonali Parbhoo 等ICML 2020 · 被引用 67 次
相关 Paper
- SOPE: Spectrum of Off-Policy EstimatorsChristina J. Yuan, Yash Chandak, Stephen Giguere, Philip S. Thomas 等NeurIPS 2021 · 被引用 6 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance SamplingYao Liu, Pierre-Luc Bacon, Emma BrunskillICML 2020 · 被引用 49 次
- Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State AbstractionBrahma S. Pavse, Josiah P. HannaAAAI 2023 · 被引用 9 次
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 被引用 26 次
