Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic Policies
Nathan Kallus, Masatoshi Uehara
Abstract
Offline reinforcement learning, wherein one uses off-policy data logged by a fixed behavior policy to evaluate and learn new policies, is crucial in applications where experimentation is limited such as medicine. We study the estimation of policy value and gradient of a deterministic policy from off-policy data when actions are continuous. Targeting deterministic policies, for which action is a deterministic function of state, is crucial since optimal policies are always deterministic (up to ties). In this setting, standard importance sampling and doubly robust estimators for policy value and gradient fail because the density ratio does not exist. To circumvent this issue, we propose several new doubly robust estimators based on different kernelization approaches. We analyze the asymptotic mean-squared error of each of these under mild rate conditions for nuisance estimators. Specifically, we demonstrate how to obtain a rate that is independent of the horizon length.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56e53493-e55f-4227-9d26-5228711d57a2Cited by top-tier papers7
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Pessimism Meets Invariance: Provably Efficient Offline Mean-Field Multi-Agent RLMinshuo Chen, Yan Li, Ethan Wang, Zhuoran Yang et al.NeurIPS 2021 · 18 citations
- Deep Jump Learning for Off-Policy Evaluation in Continuous Treatment SettingsHengrui Cai, Chengchun Shi, Rui Song, Wenbin LuNeurIPS 2021 · 18 citations
- Double Machine Learning Density Estimation for Local Treatment Effects with InstrumentsYonghan Jung, Jin Tian, Elias BareinboimNeurIPS 2021 · 15 citations
- Treatment Effect Estimation for Optimal Decision-MakingDennis Frauen, Valentyn Melnychuk, Jonas Schweisthal, Mihaela van der Schaar et al.NeurIPS 2025 · 8 citations
Builds on2
Related papers
- Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsHaanvid Lee, Jongmin Lee, Yunseon Choi, Wonseok Jeon et al.NeurIPS 2022 · 7 citations
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL PoliciesHaanvid Lee, Tri Wahyu Guntara, Jongmin Lee, Yung-Kyun Noh et al.ICLR 2024 · 3 citations
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 1 citation
- On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent SamplesMustafa O. Karabag, Ufuk TopcuAAAI 2023 · 6 citations
