Double Reinforcement Learning for Efficient and Robust Off-Policy Evaluation
Nathan Kallus, Masatoshi Uehara
摘要
Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time the semiparametric efficiency limits of OPE in Markov decision processes (MDPs), where actions, rewards, and states are memoryless. We show existing OPE estimators may fail to be efficient in this setting. We develop a new estimator based on cross-fold estimation of q-functions and marginalized density ratios, which we term double reinforcement learning (DRL). We show that DRL is efficient when both components are estimated at fourth-root rates and is also doubly robust when only one component is consistent. We investigate these properties empirically and demonstrate the performance benefits due to harnessing memorylessness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 被引用 39 次
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 被引用 60 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Semiparametrically Efficient Off-Policy Evaluation in Linear Markov Decision ProcessesChuhan Xie, Wenhao Yang, Zhihua ZhangICML 2023 · 被引用 8 次
