Efficient Multi-Policy Evaluation for Reinforcement Learning
Shuze Daniel Liu, Claire Chen, Shangtong Zhang
Abstract
To unbiasedly evaluate multiple target policies, the dominant approach among RL practitioners is to run and evaluate each target policy separately. However, this evaluation method is far from efficient because samples are not shared across policies, and running target policies to evaluate themselves is actually not optimal. In this paper, we address these two weaknesses by designing a tailored behavior policy to reduce the variance of estimators across all target policies. Theoretically, we prove that executing this behavior policy with manyfold fewer samples outperforms on-policy evaluation on every target policy under characterized conditions. Empirically, we show our estimator has a substantially lower variance compared with previous best methods and achieves state-of-the-art performance in a broad range of environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Offline Two-Player Zero-Sum Markov Games with KL RegularizationClaire Chen, Yuheng Zhang, Xinyu Liu, Zixuan Xie et al.ICML 2026 · 2 citations
- Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement LearningVagul Mahadevan, Claire Chen, Shuze D Liu, Shangtong ZhangICML 2026
- Doubly Optimal Policy Evaluation for Reinforcement LearningShuze Daniel Liu, Claire Chen, Shangtong ZhangICLR 2025
- Efficient Policy Evaluation with Safety Constraint for Reinforcement LearningClaire Chen, Shuze Daniel Liu, Shangtong ZhangICLR 2025
Builds on2
Related papers
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 7 citations
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 1 citation
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
- In-sample Actor Critic for Offline Reinforcement LearningHongchang Zhang, Yixiu Mao, Boyuan Wang, Shuncheng He et al.ICLR 2023
