Multiple-policy Evaluation via Density Estimation
Yilei Chen, Aldo Pacchiano, Ioannis Paschalidis
Abstract
We study the multiple-policy evaluation problem where we are given a set of K policies and the goal is to evaluate their performance (expected total reward over a fixed horizon) to an accuracy ϵ with probability at least 1 -δ. We propose an algorithm named CAESAR for this problem. Our approach is based on computing an approximately optimal sampling distribution and using the data sampled from it to perform the simultaneous estimation of the policy values. CAESAR has two phases. In the first phase, we produce coarse estimates of the visitation distributions of the target policies at a low order sample complexity rate that scales with Õ( 1 ϵ ). In the second phase, we approximate the optimal sampling distribution and compute the importance weighting ratios for all target policies by minimizing a step-wise quadratic loss function inspired by the DualDICE (Nachum et al., 2019) objective. Up to low order and logarithmic terms CAESAR achieves a sample complex- , where d π is the visitation distribution of policy π, µ * is the optimal sampling distribution, and H is the horizon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65120c1f-4f1e-49cd-aabf-3ca763acc570Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 161 citations
- CoinDICE: Off-Policy Confidence Interval EstimationBo Dai, Ofir Nachum, Yinlam Chow, Lihong Li et al.NeurIPS 2020 · 96 citations
- Minimax Value Interval for Off-Policy Evaluation and Policy OptimizationNan Jiang, Jiawei HuangNeurIPS 2020 · 68 citations
- Beyond Value-Function Gaps: Improved Instance-Dependent Regret Bounds for Episodic Reinforcement LearningChristoph Dann, Teodor Vanislavov Marinov, Mehryar Mohri, Julian ZimmertNeurIPS 2021 · 41 citations
- Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual BoundsYihao Feng, Ziyang Tang, Na Zhang, Qiang LiuICLR 2021 · 13 citations
Related papers
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li et al.NeurIPS 2020 · 125 citations
- Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State AbstractionBrahma S. Pavse, Josiah P. HannaAAAI 2023 · 9 citations
- SOPE: Spectrum of Off-Policy EstimatorsChristina J. Yuan, Yash Chandak, Stephen Giguere, Philip S. Thomas et al.NeurIPS 2021 · 6 citations
- Distributional Offline Policy Evaluation with Predictive Error GuaranteesRunzhe Wu, Masatoshi Uehara, Wen SunICML 2023 · 19 citations
- Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive ApproachRiccardo Poiani, Nicole Nobili, Alberto Maria Metelli, Marcello RestelliNeurIPS 2023 · 3 citations
