Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds
Yihao Feng, Ziyang Tang, Na Zhang, Qiang Liu
摘要
Off-policy evaluation (OPE) is the task of estimating the expected reward of a given policy based on offline data previously collected under different policies. Therefore, OPE is a key step in applying reinforcement learning to real-world domains such as medical treatment, where interactive data collection is expensive or even unsafe. As the observed data tends to be noisy and limited, it is essential to provide rigorous uncertainty quantification, not just a point estimation, when applying OPE to make high stakes decisions. This work considers the problem of constructing non-asymptotic confidence intervals in infinite-horizon off-policy evaluation, which remains a challenging open question. We develop a practical algorithm through a primal-dual optimization-based approach, which leverages the kernel Bellman loss (KBL) of Feng et al. (2019) and a new martingale concentration inequality of KBL applicable to timedependent data with unknown mixing conditions. Our algorithm makes minimum assumptions on the data and the function class of the Q-function, and works for the behavior-agnostic settings where the data is collected under a mix of arbitrary unknown behavior policies. We present empirical results that clearly demonstrate the advantages of our approach over existing methods. * Equal contribution. This article is an extended version of Feng et al. (2021) in ICLR 2021.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Bootstrapping Fitted Q-Evaluation for Off-Policy InferenceBotao Hao, Xiang Ji, Yaqi Duan, Hao Lu 等ICML 2021 · 被引用 46 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- Off-Policy Evaluation for Action-Dependent Non-stationary EnvironmentsYash Chandak, Shiv Shankar, Nathaniel D. Bastian, Bruno C. da Silva 等NeurIPS 2022 · 被引用 7 次
- Multiple-policy Evaluation via Density EstimationYilei Chen, Aldo Pacchiano, Ioannis PaschalidisICML 2025
它引用的顶会 Paper13
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 被引用 184 次
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 被引用 161 次
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 125 次
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 被引用 107 次
相关 Paper
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 被引用 45 次
- Sample Complexity of Nonparametric Off-Policy Evaluation on Low-Dimensional Manifolds using Deep NetworksXiang Ji, Minshuo Chen, Mengdi Wang, Tuo ZhaoICLR 2023 · 被引用 1 次
- Bellman Residual Orthogonalization for Offline Reinforcement LearningAndrea Zanette, Martin J. WainwrightNeurIPS 2022 · 被引用 14 次
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li 等NeurIPS 2020 · 被引用 96 次
- Off-Policy Interval Estimation with Lipschitz Value IterationZiyang Tang, Yihao Feng, Na Zhang, Jian Peng 等NeurIPS 2020 · 被引用 6 次
