Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic Policies
Nathan Kallus, Masatoshi Uehara
摘要
Offline reinforcement learning, wherein one uses off-policy data logged by a fixed behavior policy to evaluate and learn new policies, is crucial in applications where experimentation is limited such as medicine. We study the estimation of policy value and gradient of a deterministic policy from off-policy data when actions are continuous. Targeting deterministic policies, for which action is a deterministic function of state, is crucial since optimal policies are always deterministic (up to ties). In this setting, standard importance sampling and doubly robust estimators for policy value and gradient fail because the density ratio does not exist. To circumvent this issue, we propose several new doubly robust estimators based on different kernelization approaches. We analyze the asymptotic mean-squared error of each of these under mild rate conditions for nuisance estimators. Specifically, we demonstrate how to obtain a rate that is independent of the horizon length.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Pessimism Meets Invariance: Provably Efficient Offline Mean-Field Multi-Agent RLMinshuo Chen, Yan Li, Ethan Wang, Zhuoran Yang 等NeurIPS 2021 · 被引用 18 次
- Deep Jump Learning for Off-Policy Evaluation in Continuous Treatment SettingsHengrui Cai, Chengchun Shi, Rui Song, Wenbin LuNeurIPS 2021 · 被引用 18 次
- Double Machine Learning Density Estimation for Local Treatment Effects with InstrumentsYonghan Jung, Jin Tian, Elias BareinboimNeurIPS 2021 · 被引用 15 次
- Treatment Effect Estimation for Optimal Decision-MakingDennis Frauen, Valentyn Melnychuk, Jonas Schweisthal, Mihaela van der Schaar 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper2
相关 Paper
- Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsHaanvid Lee, Jongmin Lee, Yunseon Choi, Wonseok Jeon 等NeurIPS 2022 · 被引用 7 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL PoliciesHaanvid Lee, Tri Wahyu Guntara, Jongmin Lee, Yung-Kyun Noh 等ICLR 2024 · 被引用 3 次
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 被引用 1 次
- On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent SamplesMustafa O. Karabag, Ufuk TopcuAAAI 2023 · 被引用 6 次
