Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
Hongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang, Chengchun Shi
摘要
This paper studies off-policy evaluation (OPE) in reinforcement learning with a focus on behavior policy estimation for importance sampling. Prior work has shown empirically that estimating a history-dependent behavior policy can lead to lower mean squared error (MSE) even when the true behavior policy is Markovian. However, the question of why the use of history should lower MSE remains open. In this paper, we theoretically demystify this paradox by deriving a biasvariance decomposition of the MSE of ordinary importance sampling (IS) estimators, demonstrating that history-dependent behavior policy estimation decreases their asymptotic variances while increasing their finite-sample biases. Additionally, as the estimated behavior policy conditions on a longer history, we show a consistent decrease in variance. We extend these findings to a range of other OPE estimators, including the sequential IS estimator, the doubly robust estimator and the marginalized IS estimator, with the behavior policy estimated either parametrically or nonparametrically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia 等NeurIPS 2025 · 被引用 2 次
- Designing Time Series Experiments in A/B Testing with Transformer Reinforcement LearningXiangkun Wu, Qianglin Wen, Yingying Zhang, Hongtu Zhu 等ICLR 2026 · 被引用 1 次
- Balancing Interference and Correlation in Spatial Experimental Designs: A Causal Graph Cut ApproachJin Zhu, Jingyi Li, Hongyi Zhou, Yinan Lin 等ICML 2025
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li 等NeurIPS 2020 · 被引用 96 次
- Off-Policy Evaluation in Partially Observable EnvironmentsGuy Tennenholtz, Uri Shalit, Shie MannorAAAI 2020 · 被引用 91 次
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 被引用 81 次
相关 Paper
- SOPE: Spectrum of Off-Policy EstimatorsChristina J. Yuan, Yash Chandak, Stephen Giguere, Philip S. Thomas 等NeurIPS 2021 · 被引用 6 次
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht 等NeurIPS 2022 · 被引用 19 次
- State Relevance for Off-Policy EvaluationSimon P. Shen, Yecheng Jason Ma, Omer Gottesman, Finale Doshi-VelezICML 2021 · 被引用 6 次
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 被引用 16 次
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 被引用 26 次
