Doubly Optimal Policy Evaluation for Reinforcement Learning
Shuze Daniel Liu, Claire Chen, Shangtong Zhang
Abstract
Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting policy or data-processing method substantially deteriorates the variance of evaluation results over long time steps. Thus, policy evaluation often suffers from large variance and requires massive data to achieve the desired accuracy. In this work, we design an optimal combination of data-collecting policy and data-processing baseline. Theoretically, we prove our doubly optimal policy evaluation method is unbiased and guaranteed to have lower variance than previously best-performing methods. Empirically, compared with previous works, we show our method reduces variance substantially and achieves superior empirical performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Offline Two-Player Zero-Sum Markov Games with KL RegularizationClaire Chen, Yuheng Zhang, Xinyu Liu, Zixuan Xie et al.ICML 2026 · 2 citations
- Efficient Multi-Policy Evaluation for Reinforcement LearningShuze Daniel Liu, Claire Chen, Shangtong ZhangAAAI 2025 · 2 citations
- Designing Time Series Experiments in A/B Testing with Transformer Reinforcement LearningXiangkun Wu, Qianglin Wen, Yingying Zhang, Hongtu Zhu et al.ICLR 2026 · 1 citation
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement LearningVagul Mahadevan, Claire Chen, Shuze D Liu, Shangtong ZhangICML 2026
Builds on4
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 7 citations
- Efficient Multi-Policy Evaluation for Reinforcement LearningShuze Daniel Liu, Claire Chen, Shangtong ZhangAAAI 2025 · 2 citations
- Efficient Policy Evaluation with Safety Constraint for Reinforcement LearningClaire Chen, Shuze Daniel Liu, Shangtong ZhangICLR 2025
Related papers
- Offline RL Without Off-Policy EvaluationDavid Brandfonbrener, Will Whitney, Rajesh Ranganath, Joan BrunaNeurIPS 2021 · 217 citations
- Near-Optimal Offline Reinforcement Learning via Double Variance ReductionMing Yin, Yu Bai, Yu-Xiang WangNeurIPS 2021 · 72 citations
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 1 citation
- Beyond Variance Reduction: Understanding the True Impact of Baselines on Policy OptimizationWesley Chung, Valentin Thomas, Marlos C. Machado, Nicolas Le RouxICML 2021 · 35 citations
- Safe Exploration for Efficient Policy Evaluation and ComparisonRunzhe Wan, Branislav Kveton, Rui SongICML 2022 · 16 citations
